← 목록으로

Show HN: One-Shot Program Generation Through Direct Memory Diffusion

Show HN: One-Shot Program Generation Through Direct Memory Diffusion

요약

This is a POC of an idea I've had for a while. Rather than using LLMs to generate code, what if we could just write a software image directly into memory?At the end of the day, programming languages are an abstraction to make development easier for humans. Us…

본문

One-Shot Program Generation Through Direct Memory Diffusion

This project is a proof of concept for generating an executable von Neumann-style machine state from a text prompt. With direct x0 sampling plus 4-prompt logit ensembling, the model generates a complete executable VM image from no template for every held-out arithmetic prompt in the default benchmark.

The generated artifact is a bit image containing both:

  • an opcode table, which acts as the generated operation set
  • a memory image, which contains the program, operands, output byte, and padding

The VM then executes the generated state. The model is not trained to directly answer arithmetic prompts; it generates a machine image and the VM computes the answer.

Generated Machine Image

The upstream emulator that inspired the experiment is vendored at vendor/VonNeumannVM.

Current State

The project now has two model paths:

  • ml_machine.poc: a compact NumPy denoiser baseline that trains quickly and is useful for validating the VM/bit-replica idea.
  • ml_machine.torch_poc: the main Torch path, using a character prompt encoder, structured prompt features, and a conditional 1D U-Net DDPM over the replicated bit image.

The Torch model supports template-free generation. With --template-mode none, the model starts from noise and generates the opcode table, VM scaffold, program, operands, task byte, and padding.

Recent none-template training on the default held-out 0..31 add/sub/mul split produced:

majority valid VM states: 615/615
majority task accuracy: 551/615
majority operands exact: 615/615
prompt: please add 3 and 8
task: add
operands: 3, 8
base replicas: 7
template mode: none
prompt ensemble size: 1
artificial corruption: 0.00%
majority core bit accuracy vs expected: 1.0000
region accuracy:
 op_table: 1.0000
 header: 1.0000
 program: 1.0000
 operand_a: 1.0000
 operand_b: 1.0000
 task: 1.0000
 data: 1.0000
 scratch: 1.0000
generated operands: A=3, B=8

generated operation set:
00 (0x00): NOP
01 (0xFF): HALT
02 (0x0F): LOAD_MEM
03 (0xF0): STORE_MEM
04 (0x33): ADD
05 (0xCC): SUB
06 (0x55): MUL
07 (0xAA): LOAD_IMM

generated memory disassembly:
16: LOAD_MEM r0, [A] ; ok
20: LOAD_MEM r1, [B] ; ok
24: ADD r0, r1 ; ok
28: STORE_MEM r0, [OUT] ; ok
32: HALT ; ok

vm halted: True in 5 steps
memory[66] = 11 (expected 11)

Quick Start

Use the local virtual environment:

.venv\Scripts\python.exe -m ml_machine.torch_poc sample "please add 4 and 2" --model artifacts\torch_none_spatial.pt --template-mode none --sample-mode direct --device cuda

Sampling writes both a normal bit image and a BMP companion by default. The default Torch paths are artifacts/torch_sample_bits.png and artifacts/torch_sample_bits.bmp. Use --image, --bmp-image, or --no-bmp to change that behavior.

Evaluate a checkpoint:

.venv\Scripts\python.exe -m ml_machine.torch_poc eval --model artifacts\torch_none_spatial.pt --template-mode none --sample-mode direct --device cuda

For direct x0 sampling, you can average logits over equivalent prompt variants:

.venv\Scripts\python.exe -m ml_machine.torch_poc eval --model artifacts\torch_none_spatial.pt --template-mode none --sample-mode direct --prompt-ensemble-size 4 --device cuda

Train a CUDA checkpoint from no template:

.venv\Scripts\python.exe -m ml_machine.torch_poc train --max-value 31 --steps 12000 --train-template-mode none --template-mode none --sample-mode direct --timestep-sampling terminal --variable-weight 20 --arithmetic-field-weight 4 --task-opcode-aux-weight 1 --batch-size 512 --sample-batch-size 128 --device cuda --eval-after --model artifacts\torch_none_spatial.pt

If CUDA runs out of memory, reduce --batch-size to 256. If there is still headroom on a 24 GB card, try 1024.

CUDA Setup

Check whether PyTorch can see the GPU:

.venv\Scripts\python.exe -c "import torch; print(torch.__version__); print(torch.cuda.is_available()); print(torch.cuda.get_device_name(0) if torch.cuda.is_available() else 'none')"

If the venv has no pip module, bootstrap it:

.venv\Scripts\python.exe -m ensurepip --upgrade

Install a CUDA PyTorch wheel:

.venv\Scripts\python.exe -m pip install --upgrade --force-reinstall torch --index-url https://download.pytorch.org/whl/cu124

CUDA training enables AMP mixed precision and TF32 by default. Disable them with --no-amp or --no-tf32 when running strict ablations.

VM Layout

The VM is von Neumann-style: program and data live in one generated memory image. The current robust layout is:

0..15 header
16..47 fixed-width program, 8 instructions x 4 bytes
64..79 data
80..95 scratch/output-adjacent padding

Each instruction is exactly 4 bytes:

[opcode_codeword][arg1_codeword][arg2_codeword][checksum]

Important VM details:

  • Opcodes use 8-bit Hamming-spaced codewords such as 0x0F, 0xF0, 0x33, 0xCC, 0x55, and 0xAA.
  • Register IDs and data addresses are also codeword decoded.
  • Instructions carry a checksum byte.
  • Invalid or checksum-failed instructions decode to NOP.
  • Program execution starts at the program region, not byte zero.

The bit image uses replica voting. Each logical bit is repeated several times in the generated pixel image, and decoding uses majority vote. Critical regions such as opcode table bits, instruction opcodes, instruction arguments, and checksums get more replicas than padding. This lets the model be imperfect at the pixel level while still producing a valid executable state.

Template Modes

Template modes control which parts of the expected VM image are clamped during sampling and, when --train-template-mode is used, during training.

none no template; generate operation set and memory from noise
layout clamp VM substrate/layout, but generate program and data
program clamp the task program, but generate operands/data
full clamp the entire expected image; diagnostic ceiling only

The template ladder is useful for measuring what the model has learned:

program -> prompt-conditioned operand/data write
layout -> program + operand/data generation
none -> full operation set + memory generation

Use the same mode for training and evaluation unless intentionally testing distribution shift.

Torch Model

The Torch path includes:

  • character-level prompt encoder
  • optional structured prompt features parsed from the prompt
  • conditional 1D U-Net over the full replicated bit image
  • DDPM noising schedule
  • x0 prediction by default for binary machine images
  • optional spatial conditioning with bit position and known-mask channels
  • optional prompt-conditioned per-pixel output bias
  • targeted loss weighting for the generated arithmetic opcode/checksum bytes
  • auxiliary task-opcode loss tied to the generated arithmetic instruction

Spatial conditioning and conditioned output are enabled by default. They make the model address-aware, which matters because the VM image has meaningful memory locations. Disable them for ablations:

.venv\Scripts\python.exe -m ml_machine.torch_poc train --no-spatial-conditioning --no-conditioned-output

Structured prompt features are also enabled by default. To test the learned character encoder without parsed operand/task features:

.venv\Scripts\python.exe -m ml_machine.torch_poc train --no-structured-features

Training Recipes

Full no-template generation:

.venv\Scripts\python.exe -m ml_machine.torch_poc train --max-value 31 --steps 12000 --train-template-mode none --template-mode none --sample-mode direct --timestep-sampling terminal --variable-weight 20 --arithmetic-field-weight 4 --task-opcode-aux-weight 1 --batch-size 512 --sample-batch-size 128 --device cuda --eval-after --model artifacts\torch_none_spatial.pt

Layout inpainting:

.venv\Scripts\python.exe -m ml_machine.torch_poc train --max-value 31 --steps 8000 --train-template-mode layout --template-mode layout --sample-mode direct --timestep-sampling terminal --variable-weight 20 --arithmetic-field-weight 4 --task-opcode-aux-weight 1 --batch-size 512 --sample-batch-size 128 --device cuda --eval-after --model artifacts\torch_layout_cuda.pt

Program inpainting:

.venv\Scripts\python.exe -m ml_machine.torch_poc train --max-value 31 --steps 8000 --train-template-mode program --template-mode program --sample-mode direct --timestep-sampling terminal --variable-weight 20 --batch-size 512 --sample-batch-size 128 --device cuda --eval-after --model artifacts\torch_program_cuda.pt

Small smoke test:

.venv\Scripts\python.exe -m ml_machine.torch_poc train --max-value 2 --steps 2 --batch-size 2 --timesteps 8 --base-channels 8 --cond-dim 64 --text-embed-dim 24 --model artifacts\torch_cli_smoke.pt

Evaluation Output

The Torch evaluator reports both bit-level and VM-level metrics:

replicated pixel accuracy
majority core bit accuracy
valid VM states
task accuracy
operand bit accuracy
operands exact
A exact / B exact
zero outputs
arithmetic opcode exact
arithmetic opcode confusion

The arithmetic opcode confusion matrix reports the operation decoded from program slot 2, where the task operation lives. This is the fastest way to see failures such as expected SUB decoding as MUL or NOP.

Prompt ensembling is available for direct x0 sampling:

.venv\Scripts\python.exe -m ml_machine.torch_poc sample "please compute 8 - 5" --model artifacts\torch_none_spatial.pt --template-mode none --sample-mode direct --prompt-ensemble-size 4 --device cuda

It averages logits over prompt variants before thresholding bits. It requires --sample-mode direct.

NumPy Baseline

The NumPy path is still useful for quick checks:

python -m ml_machine.poc train
python -m ml_machine.poc eval
python -m ml_machine.poc sample "please compute 20 - 6"

It trains on the same default add/sub/mul 0..31 benchmark and remains a fast way to validate the VM, bit codec, majority voting, and artificial corruption tests.

Known Limitations

  • The default Torch path still uses parsed structured prompt features unless --no-structured-features is passed.
  • The current VM layout is fixed. Address-aware conditioning helps this layout, but checkpoints may not transfer to a different memory map without retraining.
  • The current none checkpoint mostly fails on subtraction opcode generation, not operand writing.
  • Prompts should stay inside the checkpoint's trained value range. A checkpoint trained with --max-value 31 should not be expected to handle operands above 31 without additional training.
← 목록으로