Series

OKReport Local AI Build Lab

Same brief. Local models. Real implementation. A public benchmark series. Each benchmark gives local coding models the same feature brief, checks their work against a real private codebase, and publishes what a reviewer could reproduce—including the parts that failed.

How the lab runs

Every benchmark uses the same publishing rules, so successes and failures stay comparable.

  1. One brief, no per-model tuning

    Every implementation model receives the same written brief, the same starting codebase, and the same definition of finished.

  2. Outcomes checked before publication

    The reviewer reruns the available backend, frontend, browser, artifact, and repository checks instead of repeating what a model claimed.

  3. Sanitized by default

    Client names, internal domains, addresses, customer data, credentials, and proprietary source never leave the private repository.

What we compare

No composite score hides the important distinction between working code, safe integration, and reproducible proof.

  • Functional result

    What the model completed and whether the real user flow could be reached.

  • Safety and architecture

    Whether it preserved existing permissions, confirmation boundaries, and system-of-record paths.

  • Verification

    Which backend, frontend, browser, artifact, and repository checks the reviewer could reproduce.

  • Autonomy

    What human help the run received, separated by the kind of unblock rather than collapsed into one number.

  • Interface evidence

    An authentic runnable capture when one exists, and an explicit unavailable state when it does not.

Outcome statuses

Reviewer-reproduced artifact flow
Read the model record for the exact evidence boundary behind this status.
Complete with a verification gap
Read the model record for the exact evidence boundary behind this status.
Partial
Read the model record for the exact evidence boundary behind this status.
Incomplete
Read the model record for the exact evidence boundary behind this status.

How to read a benchmark

Three things to notice before interpreting one run as a general model claim.

  1. Read the status before the implementation detail

    Complete, verification gap, partial, and incomplete describe different evidence boundaries; they are not a numerical ranking.

  2. Treat missing proof as a gap

    An unavailable check is reported as unavailable. It is never inferred as a pass from nearby green tests.

  3. Read the human help in context

    Completion nudges, infrastructure access, permission unblocks, and technical steering are different interventions and are reported separately.

Benchmarks

1 benchmark published. Each one includes the brief, outcomes, human help, interface evidence, and verification notes.

Benchmark 1 · September 2026

Reviewer-verified outcomes

Integrating OpenWebUI with ERPNext

Six locally served coding models implemented the same ERPNext assistant feature. We compared what they built, what worked, what broke, and how much help each needed.

Models compared
6
Runnable UI captures
2
Shared brief
One
Result format
Right, wrong, proof
  • Qwen 3.8 Flash Next FP8Live chain verified; operator access blocked
  • MiMo V2.6 FlashReviewer-reproduced after configuration repair
  • DeepSeek V4 Flash 0731Incomplete at the fair cutoff
  • Qwen 3.8 Flash Next NVFP4Partial: chat bridge missing
  • Qwen 3.8 27B BF16Incomplete
  • Gemma 4 31B BF16Incomplete