Code generation with open-weight models, fully offline

Building an offline harness for open-weight models: quantization trade-offs and what small models do well.

Published

Open-weight modelsTooling

Why run models locally

Open-weight models can be downloaded and run on your own hardware. Nothing about your codebase is sent anywhere, there is no per-request cost, it works on a plane, and the model does not change underneath you — the same weights give the same behaviour next month.

Quantization trade-offs

Quantization stores weights at lower precision so a model fits in less memory and runs faster. Lower precision saves more memory but can reduce output quality, especially on tasks that need precise syntax. The practical approach is to try a few levels on your own real tasks and pick the smallest one whose output you would still accept after review.

High context, low reasoning

Smaller models are weaker at open-ended reasoning but good at pattern-following when the pattern is in front of them. So give them the context instead of asking them to infer it: the relevant types, the schema, an example of similar code, and a precise description of the change.

Break big changes into small, checkable steps. A modest local model that completes ten narrow steps correctly is more useful than one that attempts the whole change and gets it half right.

  • Put schemas and examples in the prompt
  • Ask for one change at a time
  • Review and test every output before keeping it
  • Save the bench config next to the result it produced

Related projects