요약
This is a POC of an idea I've had for a while. Rather than using LLMs to generate code, what if we could just write a software image directly into memory?At the end of the day, programming languages are an abstraction to make development easier for humans. Us…
본문
One-Shot Program Generation Through Direct Memory DiffusionThis project is a proof of concept for generating an executable von Neumann-style machine state from a text prompt. With direct x0 sampling plus 4-prompt logit ensembling, the model generates a complete executable VM image from no template for every held-out arithmetic prompt in the default benchmark.
The generated artifact is a bit image containing both:
- an opcode table, which acts as the generated operation set
- a memory image, which contains the program, operands, output byte, and padding
The VM then executes the generated state. The model is not trained to directly answer arithmetic prompts; it generates a machine image and the VM computes the answer.
The upstream emulator that inspired the experiment is vendored at
vendor/VonNeumannVM.
Current State
The project now has two model paths:
ml_machine.poc: a compact NumPy denoiser baseline that trains quickly and is useful for validating the VM/bit-replica idea.ml_machine.torch_poc: the main Torch path, using a character prompt encoder, structured prompt features, and a conditional 1D U-Net DDPM over the replicated bit image.
The Torch model supports template-free generation. With --template-mode none,
the model starts from noise and generates the opcode table, VM scaffold, program,
operands, task byte, and padding.
Recent none-template training on the default held-out 0..31 add/sub/mul split
produced:
majority valid VM states: 615/615
majority task accuracy: 551/615
majority operands exact: 615/615
prompt: please add 3 and 8
task: add
operands: 3, 8
base replicas: 7
template mode: none
prompt ensemble size: 1
artificial corruption: 0.00%
majority core bit accuracy vs expected: 1.0000
region accuracy:
op_table: 1.0000
header: 1.0000
program: 1.0000
operand_a: 1.0000
operand_b: 1.0000
task: 1.0000
data: 1.0000
scratch: 1.0000
generated operands: A=3, B=8
generated operation set:
00 (0x00): NOP
01 (0xFF): HALT
02 (0x0F): LOAD_MEM
03 (0xF0): STORE_MEM
04 (0x33): ADD
05 (0xCC): SUB
06 (0x55): MUL
07 (0xAA): LOAD_IMM
generated memory disassembly:
16: LOAD_MEM r0, [A] ; ok
20: LOAD_MEM r1, [B] ; ok
24: ADD r0, r1 ; ok
28: STORE_MEM r0, [OUT] ; ok
32: HALT ; ok
vm halted: True in 5 steps
memory[66] = 11 (expected 11)
Quick Start
Use the local virtual environment:
.venv\Scripts\python.exe -m ml_machine.torch_poc sample "please add 4 and 2" --model artifacts\torch_none_spatial.pt --template-mode none --sample-mode direct --device cuda
Sampling writes both a normal bit image and a BMP companion by default. The default
Torch paths are artifacts/torch_sample_bits.png and
artifacts/torch_sample_bits.bmp. Use --image, --bmp-image, or --no-bmp to
change that behavior.
Evaluate a checkpoint:
.venv\Scripts\python.exe -m ml_machine.torch_poc eval --model artifacts\torch_none_spatial.pt --template-mode none --sample-mode direct --device cuda
For direct x0 sampling, you can average logits over equivalent prompt variants:
.venv\Scripts\python.exe -m ml_machine.torch_poc eval --model artifacts\torch_none_spatial.pt --template-mode none --sample-mode direct --prompt-ensemble-size 4 --device cuda
Train a CUDA checkpoint from no template:
.venv\Scripts\python.exe -m ml_machine.torch_poc train --max-value 31 --steps 12000 --train-template-mode none --template-mode none --sample-mode direct --timestep-sampling terminal --variable-weight 20 --arithmetic-field-weight 4 --task-opcode-aux-weight 1 --batch-size 512 --sample-batch-size 128 --device cuda --eval-after --model artifacts\torch_none_spatial.pt
If CUDA runs out of memory, reduce --batch-size to 256. If there is still
headroom on a 24 GB card, try 1024.
CUDA Setup
Check whether PyTorch can see the GPU:
.venv\Scripts\python.exe -c "import torch; print(torch.__version__); print(torch.cuda.is_available()); print(torch.cuda.get_device_name(0) if torch.cuda.is_available() else 'none')"
If the venv has no pip module, bootstrap it:
.venv\Scripts\python.exe -m ensurepip --upgrade
Install a CUDA PyTorch wheel:
.venv\Scripts\python.exe -m pip install --upgrade --force-reinstall torch --index-url https://download.pytorch.org/whl/cu124
CUDA training enables AMP mixed precision and TF32 by default. Disable them with
--no-amp or --no-tf32 when running strict ablations.
VM Layout
The VM is von Neumann-style: program and data live in one generated memory image. The current robust layout is:
0..15 header
16..47 fixed-width program, 8 instructions x 4 bytes
64..79 data
80..95 scratch/output-adjacent padding
Each instruction is exactly 4 bytes:
[opcode_codeword][arg1_codeword][arg2_codeword][checksum]
Important VM details:
- Opcodes use 8-bit Hamming-spaced codewords such as
0x0F,0xF0,0x33,0xCC,0x55, and0xAA. - Register IDs and data addresses are also codeword decoded.
- Instructions carry a checksum byte.
- Invalid or checksum-failed instructions decode to
NOP. - Program execution starts at the program region, not byte zero.
The bit image uses replica voting. Each logical bit is repeated several times in the generated pixel image, and decoding uses majority vote. Critical regions such as opcode table bits, instruction opcodes, instruction arguments, and checksums get more replicas than padding. This lets the model be imperfect at the pixel level while still producing a valid executable state.
Template Modes
Template modes control which parts of the expected VM image are clamped during
sampling and, when --train-template-mode is used, during training.
none no template; generate operation set and memory from noise
layout clamp VM substrate/layout, but generate program and data
program clamp the task program, but generate operands/data
full clamp the entire expected image; diagnostic ceiling only
The template ladder is useful for measuring what the model has learned:
program -> prompt-conditioned operand/data write
layout -> program + operand/data generation
none -> full operation set + memory generation
Use the same mode for training and evaluation unless intentionally testing distribution shift.
Torch Model
The Torch path includes:
- character-level prompt encoder
- optional structured prompt features parsed from the prompt
- conditional 1D U-Net over the full replicated bit image
- DDPM noising schedule
x0prediction by default for binary machine images- optional spatial conditioning with bit position and known-mask channels
- optional prompt-conditioned per-pixel output bias
- targeted loss weighting for the generated arithmetic opcode/checksum bytes
- auxiliary task-opcode loss tied to the generated arithmetic instruction
Spatial conditioning and conditioned output are enabled by default. They make the model address-aware, which matters because the VM image has meaningful memory locations. Disable them for ablations:
.venv\Scripts\python.exe -m ml_machine.torch_poc train --no-spatial-conditioning --no-conditioned-output
Structured prompt features are also enabled by default. To test the learned character encoder without parsed operand/task features:
.venv\Scripts\python.exe -m ml_machine.torch_poc train --no-structured-features
Training Recipes
Full no-template generation:
.venv\Scripts\python.exe -m ml_machine.torch_poc train --max-value 31 --steps 12000 --train-template-mode none --template-mode none --sample-mode direct --timestep-sampling terminal --variable-weight 20 --arithmetic-field-weight 4 --task-opcode-aux-weight 1 --batch-size 512 --sample-batch-size 128 --device cuda --eval-after --model artifacts\torch_none_spatial.pt
Layout inpainting:
.venv\Scripts\python.exe -m ml_machine.torch_poc train --max-value 31 --steps 8000 --train-template-mode layout --template-mode layout --sample-mode direct --timestep-sampling terminal --variable-weight 20 --arithmetic-field-weight 4 --task-opcode-aux-weight 1 --batch-size 512 --sample-batch-size 128 --device cuda --eval-after --model artifacts\torch_layout_cuda.pt
Program inpainting:
.venv\Scripts\python.exe -m ml_machine.torch_poc train --max-value 31 --steps 8000 --train-template-mode program --template-mode program --sample-mode direct --timestep-sampling terminal --variable-weight 20 --batch-size 512 --sample-batch-size 128 --device cuda --eval-after --model artifacts\torch_program_cuda.pt
Small smoke test:
.venv\Scripts\python.exe -m ml_machine.torch_poc train --max-value 2 --steps 2 --batch-size 2 --timesteps 8 --base-channels 8 --cond-dim 64 --text-embed-dim 24 --model artifacts\torch_cli_smoke.pt
Evaluation Output
The Torch evaluator reports both bit-level and VM-level metrics:
replicated pixel accuracy
majority core bit accuracy
valid VM states
task accuracy
operand bit accuracy
operands exact
A exact / B exact
zero outputs
arithmetic opcode exact
arithmetic opcode confusion
The arithmetic opcode confusion matrix reports the operation decoded from program
slot 2, where the task operation lives. This is the fastest way to see failures
such as expected SUB decoding as MUL or NOP.
Prompt ensembling is available for direct x0 sampling:
.venv\Scripts\python.exe -m ml_machine.torch_poc sample "please compute 8 - 5" --model artifacts\torch_none_spatial.pt --template-mode none --sample-mode direct --prompt-ensemble-size 4 --device cuda
It averages logits over prompt variants before thresholding bits. It requires
--sample-mode direct.
NumPy Baseline
The NumPy path is still useful for quick checks:
python -m ml_machine.poc train python -m ml_machine.poc eval python -m ml_machine.poc sample "please compute 20 - 6"
It trains on the same default add/sub/mul 0..31 benchmark and remains a fast way
to validate the VM, bit codec, majority voting, and artificial corruption tests.
Known Limitations
- The default Torch path still uses parsed structured prompt features unless
--no-structured-featuresis passed. - The current VM layout is fixed. Address-aware conditioning helps this layout, but checkpoints may not transfer to a different memory map without retraining.
- The current
nonecheckpoint mostly fails on subtraction opcode generation, not operand writing. - Prompts should stay inside the checkpoint's trained value range. A checkpoint
trained with
--max-value 31should not be expected to handle operands above31without additional training.
