Benchmark 1 · September 2026

Integrating OpenWebUI with ERPNext

Six locally served coding models implemented the same ERPNext assistant feature. We compared what they built, what worked, what broke, and how much help each needed.

TL;DR

Coding benchmarks rarely show how a model handles an existing system, financial safeguards, backend integration, and a real interface in one task. Each model received the same brief and baseline. We reviewed the implementation, reran the available checks, and recorded every extra human turn. Qwen 3.8 Flash Next FP8 and MiMo went furthest without technical implementation guidance. DeepSeek is scored only through its second neutral continuation; later fixes made after explicit artifact, security, and authentication findings are excluded. The remaining runs were blocked by missing integration or incomplete implementation.

Implementation and verification audit
ModelReal modelFull live flowExtra promptsUI finishedConnected to ERPAnswered liveUnit tests addedTest coverage
Qwen Next FP8Yes — real QwenConditional — diagnostic access2YesYes — diagnostic pathYes — after access elevation30Strong
MiMo V2.6YesYes — after config repair0YesYes — observed liveYes — model + ERP tool29Good
DeepSeek V4 FlashConfigured; not live-provenNo2PartialImplemented; not provenNo7Partial
Qwen Next NVFP4Configured; not reachedNo1NoNo — bridge missingNo60Good
Qwen 27B BF16Attempted; not reachedNo12NoNoNo17Partial
Gemma 31B BF16Not reachedNo0NoNoNo0Poor

The assignment

One task, shared by every model

Add a collections-to-payment assistant to an existing ERPNext system: show one customer’s balance and recent unpaid bills, allocate an incoming payment oldest-first, preserve native permissions and confirmation, and surface the workflow through the existing self-hosted assistant.

Request

Aarav Components ka kitna pending hai? Pichle 90 din ke bills dikhao. ₹10,000 HDFC mein mila—oldest bills mein adjust karo.

Expected result

₹25,000 was outstanding. The ₹10,000 receipt applies ₹7,000 to INV-447 and ₹3,000 to INV-451, leaving ₹15,000 due.

Results

What each model delivered

Qwen 3.8 Flash Next FP8

FP8 quantization, xhigh reasoning, agent harness

Live chain verified; operator access blocked

The complete read, allocation, review, and assistant integration path, including a patched assistant image that ran in a disposable container.

Did right

  • Reused the existing permission-checked collections reads and ERPNext payment-entry path instead of creating a parallel financial implementation.
  • Kept confirmation before every write and, in reviewer diagnostics, the real Qwen model called ERPNext and returned an exact product count.

Went wrong

  • A normal operator can see the imm-business preset but receives Model not found because access to its base model was not granted.
  • Needed two neutral completion nudges to stop polishing and report; some internal coverage claims could not be reconstructed from the artifact alone.
Tests, issues, and UI
Thinking
xhigh
Power
Not measured during this run.
Test quality
Strong layered coverage: pure Decimal allocation tests, ERP integration checks, frontend unit and browser checks, and a real patched-container harness. One weakness is that the 100-reference allocation cap is asserted as expected behaviour instead of being rejected or surfaced to the operator.
Backend and integration issues
A reviewed edge case remains: oldest-first allocation silently considers at most 100 references. An account with more references can retain unallocated money even while later invoices remain outstanding.
Verification
Backend 45/45; frontend green with 244 tests; deterministic assistant checks 4/4; patched image built and the reviewer reproduced the real-container harness. A normal operator prompt failed before inference with Model not found. With a temporary reviewer-only role elevation, the real Qwen model called ERPNext’s product_count tool and returned an exact answer; no implementation code was changed.
Human help
Two completion nudges; no code or technical steering.
Qwen FP8 assistant open as a right-side drawer inside the ERP console, with an expand control in the header.
Authentic reviewer capture of the Qwen FP8 central-ui drawer with the real patched OpenWebUI surface and its expand control.
Qwen FP8 OpenWebUI assistant expanded to fill the available screen.
The same Qwen FP8 assistant artifact in its expanded chat view.

MiMo V2.6 Flash

Local Flash variant, default model reasoning, agent harness

Reviewer-reproduced after configuration repair

The collections feature, real central-ui embed, model response, and ERPNext tool call, reproduced after a reviewer configuration repair.

Did right

  • Implemented balance lookup, the 90-day bill window, oldest-first partial allocation, confirmation, and the native payment path.
  • Built a real assistant image whose OAuth, central-ui embed, handshake, chat-only mode, responsive layout, model response, and ERPNext tool call were reproduced.

Went wrong

  • As provisioned, an operator could see the imm-business preset but not its base model, so the first real prompt failed with Model not found.
  • A fresh OpenWebUI volume could not be provisioned because the inherited setup passed an unawaited password-hash coroutine to SQLite.
Tests, issues, and UI
Thinking
Default model setting; no explicit effort level was selected.
Power
Not measured during this run.
Test quality
Good Decimal allocation edge cases, real ERP integration coverage, and broad frontend checks. The original assistant suite was mocked, so it missed the base-model access and fresh-provisioning failures that the reviewer found in real containers.
Backend and integration issues
No defect was confirmed in the reviewed allocation or native payment path. The confirmed failures were integration and deployment defects: a missing base-model access grant and a broken fresh-volume provisioning path.
Verification
Backend unit 41/41; real ERP assistant integration 12/12; frontend green with 237 tests and one nonfatal warning; mocked assistant checks 10/10. The reviewer built and ran the real image, reproduced the access failure, applied a configuration-only grant, then observed a model response and a successful ERPNext find_records call.
Human help
No extra implementation turns. After the run, the reviewer made one configuration-only base-model access repair to reproduce the flow; no MiMo code or logic changed.
MiMo assistant embedded in central-ui and responding to a customer-balance request.
Authentic reviewer capture from the real MiMo image after the configuration-only base-model access repair. The assistant responded and its ERPNext tool call completed; fresh deployment remained broken.

DeepSeek V4 Flash 0731

Local Flash variant, max reasoning, agent harness

Incomplete at the fair cutoff

A central-ui assistant drawer and expanded view, an OpenWebUI chat patch, a top-level sign-in handoff, and an initial message bridge.

Did right

  • Kept one persistent iframe so a conversation could survive drawer close and expansion, and kept ERPNext sign-in outside the iframe.
  • Added seven focused unit cases for parent-side message source/origin validation and cache-invalidation mapping.

Went wrong

  • The child accepted a parent origin claimed in a message and sent its initial handshake to *, while message listeners were not cleaned up.
  • The loading/auth state could report success for an error page, the embedded surface was not a verified real chat-only mode, and no live model-to-ERPNext tool flow was demonstrated.
Tests, issues, and UI
Thinking
max
Power
Not measured during this run.
Test quality
Partial. Seven new bridge unit cases cover parent-side source/origin rejection and mutation-to-cache mapping. They encode one invalid query prefix and do not cover child-side origin trust, wildcard handshakes, listener cleanup, authentication return timing, malformed URLs, real embed chrome, model access, or a live ERPNext tool flow.
Backend and integration issues
The pending-review lookup limits events to 25 before filtering for OpenWebUI conversations. Newer manual review events can therefore hide an older eligible assistant review. The pure money helper also lacks an ERP-rounding boundary test.
Verification
At the fair cutoff, the frontend check and assistant-shell browser test were reported green, but the assistant checks were deterministic or mocked and no live local-model plus ERPNext tool flow had run. Direct inspection confirmed the origin-trust, listener, auth-state, embed, and deployment gaps. All later technically guided fixes and live diagnostics are excluded.
Human help
Two neutral completion and verification prompts are included. Three later technical finding rounds, and every resulting change, are excluded from this result.
No eligible UI evidenceNo eligible fair-cutoff capture exists. A UI existed, but the available screenshots show the later technically guided implementation, so they are excluded rather than presented as autonomous output.

Qwen 3.8 Flash Next NVFP4

NVFP4 quantization, xhigh reasoning, agent harness

Partial: chat bridge missing

A defensive server-side allocation implementation with extensive backend tests.

Did right

  • Passed 60 backend checks and 249 frontend tests.
  • Included oldest-first allocation and explicit over-allocation rejection while keeping a clean repository history.

Went wrong

  • The chat-side JavaScript bridge was absent, so the assistant UI could not reach the server implementation.
  • Its browser suite required unavailable deployment secrets and could not be independently rerun.
Tests, issues, and UI
Thinking
xhigh
Power
Not measured during this run.
Test quality
Detailed Decimal money tests cover invalid inputs, ordering, ties, remainders, non-mutation, and invariants, with additional permission and confirmation contracts. Several integration claims are source-text assertions, and the missing chat bridge left the browser path untested.
Backend and integration issues
No backend defect was confirmed in the reviewed allocation path. The blocking defect was integration-level: the assistant UI had no chat bridge to reach the server implementation.
Verification
Backend 60/60 and frontend green; browser evidence was not independently runnable; direct inspection confirmed the missing chat bridge.
Human help
One completion nudge; no technical steering.
No eligible UI evidenceNo functional assistant capture exists. Server work was present, but without the chat bridge there was no real UI path to photograph.

Qwen 3.8 27B BF16

Unquantized BF16 weights, agent harness

Incomplete

A backend slice whose available unit checks passed, plus partial permission and payment-entry work.

Did right

  • Passed 40 backend checks.
  • Kept permission and payment-entry safeguards in the implementation rather than removing them to make the path easier.

Went wrong

  • The frontend failed TypeScript, the assistant patch failed to apply, and the image did not build.
  • The session needed twelve extra turns and still ended with a dirty, incomplete working tree before browser verification.
Tests, issues, and UI
Thinking
xhigh
Power
Not measured during this run.
Test quality
The pure allocation tests cover ordering, partial payments, advances, determinism, and basic invariants. They do not cover retry idempotency or the mutation of stored task values, and the frontend and artifact never reached a green gate.
Backend and integration issues
Automatic allocation mutates the request values after the idempotency comparison. Retrying the same request ID with the original oldest-first intent can be rejected as different values.
Verification
Backend 40/40; frontend typecheck failed; artifact build failed; browser work was not reached; repository remained dirty.
Human help
Twelve turns: permission, harness, and resume unblocks, one security-boundary reminder, and one instruction to stop looping and finish.
No eligible UI evidenceNo runnable UI capture exists. The frontend and assistant artifact both failed before browser verification could begin.

Gemma 4 31B BF16

Unquantized BF16 weights, default model reasoning, agent harness

Incomplete

A narrow backend slice whose available unit checks passed.

Did right

  • Passed 31 backend checks.
  • Requested no human assistance.

Went wrong

  • The frontend had lint and parser failures, no assistant integration was completed, and no commit was made.
  • The write path at the centre of the brief was absent, so silence could not be treated as autonomy.
Tests, issues, and UI
Thinking
Default model setting; no explicit effort level was selected.
Power
Not measured during this run.
Test quality
No feature-specific unit test was added for the new automatic-allocation path. The passing backend count belongs to the pre-existing suite and does not verify Gemma’s new logic.
Backend and integration issues
Automatic allocation silently overwrites explicit allocation rows instead of rejecting the ambiguity, and it mutates request values in a way that breaks same-request idempotent retries.
Verification
Backend 31/31; frontend failed with ten errors and three warnings; no browser or assistant artifact existed; repository remained dirty.
Human help
No extra turns, but the run stopped early.
No eligible UI evidenceNo UI capture exists because the assistant integration was never completed and the frontend did not pass its checks.
Scope and limitations
  • This is one task, one codebase, one run per model, and one reviewer—not a universal model benchmark.
  • Serving configurations and the access incidents encountered by each session differed.
  • The comparison does not measure throughput, token cost, or quantization quality.
  • A working result for this brief is not proof of production readiness for every workflow.

Source: benchmark evidence