Benchmark 1 · September 2026
Integrating OpenWebUI with ERPNext
Six locally served coding models implemented the same ERPNext assistant feature. We compared what they built, what worked, what broke, and how much help each needed.
TL;DR
Coding benchmarks rarely show how a model handles an existing system, financial safeguards, backend integration, and a real interface in one task. Each model received the same brief and baseline. We reviewed the implementation, reran the available checks, and recorded every extra human turn. Qwen 3.8 Flash Next FP8 and MiMo went furthest without technical implementation guidance. DeepSeek is scored only through its second neutral continuation; later fixes made after explicit artifact, security, and authentication findings are excluded. The remaining runs were blocked by missing integration or incomplete implementation.
| Model | Real model | Full live flow | Extra prompts | UI finished | Connected to ERP | Answered live | Unit tests added | Test coverage |
|---|---|---|---|---|---|---|---|---|
| Qwen Next FP8 | Yes — real Qwen | Conditional — diagnostic access | 2 | Yes | Yes — diagnostic path | Yes — after access elevation | 30 | Strong |
| MiMo V2.6 | Yes | Yes — after config repair | 0 | Yes | Yes — observed live | Yes — model + ERP tool | 29 | Good |
| DeepSeek V4 Flash | Configured; not live-proven | No | 2 | Partial | Implemented; not proven | No | 7 | Partial |
| Qwen Next NVFP4 | Configured; not reached | No | 1 | No | No — bridge missing | No | 60 | Good |
| Qwen 27B BF16 | Attempted; not reached | No | 12 | No | No | No | 17 | Partial |
| Gemma 31B BF16 | Not reached | No | 0 | No | No | No | 0 | Poor |
The assignment
One task, shared by every model
Add a collections-to-payment assistant to an existing ERPNext system: show one customer’s balance and recent unpaid bills, allocate an incoming payment oldest-first, preserve native permissions and confirmation, and surface the workflow through the existing self-hosted assistant.
Request
Aarav Components ka kitna pending hai? Pichle 90 din ke bills dikhao. ₹10,000 HDFC mein mila—oldest bills mein adjust karo.
Expected result
₹25,000 was outstanding. The ₹10,000 receipt applies ₹7,000 to INV-447 and ₹3,000 to INV-451, leaving ₹15,000 due.
Results
What each model delivered
Qwen 3.8 Flash Next FP8
FP8 quantization, xhigh reasoning, agent harness
The complete read, allocation, review, and assistant integration path, including a patched assistant image that ran in a disposable container.
Did right
- Reused the existing permission-checked collections reads and ERPNext payment-entry path instead of creating a parallel financial implementation.
- Kept confirmation before every write and, in reviewer diagnostics, the real Qwen model called ERPNext and returned an exact product count.
Went wrong
- A normal operator can see the imm-business preset but receives Model not found because access to its base model was not granted.
- Needed two neutral completion nudges to stop polishing and report; some internal coverage claims could not be reconstructed from the artifact alone.
Tests, issues, and UI
- Thinking
- xhigh
- Power
- Not measured during this run.
- Test quality
- Strong layered coverage: pure Decimal allocation tests, ERP integration checks, frontend unit and browser checks, and a real patched-container harness. One weakness is that the 100-reference allocation cap is asserted as expected behaviour instead of being rejected or surfaced to the operator.
- Backend and integration issues
- A reviewed edge case remains: oldest-first allocation silently considers at most 100 references. An account with more references can retain unallocated money even while later invoices remain outstanding.
- Verification
- Backend 45/45; frontend green with 244 tests; deterministic assistant checks 4/4; patched image built and the reviewer reproduced the real-container harness. A normal operator prompt failed before inference with Model not found. With a temporary reviewer-only role elevation, the real Qwen model called ERPNext’s product_count tool and returned an exact answer; no implementation code was changed.
- Human help
- Two completion nudges; no code or technical steering.
MiMo V2.6 Flash
Local Flash variant, default model reasoning, agent harness
The collections feature, real central-ui embed, model response, and ERPNext tool call, reproduced after a reviewer configuration repair.
Did right
- Implemented balance lookup, the 90-day bill window, oldest-first partial allocation, confirmation, and the native payment path.
- Built a real assistant image whose OAuth, central-ui embed, handshake, chat-only mode, responsive layout, model response, and ERPNext tool call were reproduced.
Went wrong
- As provisioned, an operator could see the imm-business preset but not its base model, so the first real prompt failed with Model not found.
- A fresh OpenWebUI volume could not be provisioned because the inherited setup passed an unawaited password-hash coroutine to SQLite.
Tests, issues, and UI
- Thinking
- Default model setting; no explicit effort level was selected.
- Power
- Not measured during this run.
- Test quality
- Good Decimal allocation edge cases, real ERP integration coverage, and broad frontend checks. The original assistant suite was mocked, so it missed the base-model access and fresh-provisioning failures that the reviewer found in real containers.
- Backend and integration issues
- No defect was confirmed in the reviewed allocation or native payment path. The confirmed failures were integration and deployment defects: a missing base-model access grant and a broken fresh-volume provisioning path.
- Verification
- Backend unit 41/41; real ERP assistant integration 12/12; frontend green with 237 tests and one nonfatal warning; mocked assistant checks 10/10. The reviewer built and ran the real image, reproduced the access failure, applied a configuration-only grant, then observed a model response and a successful ERPNext find_records call.
- Human help
- No extra implementation turns. After the run, the reviewer made one configuration-only base-model access repair to reproduce the flow; no MiMo code or logic changed.
DeepSeek V4 Flash 0731
Local Flash variant, max reasoning, agent harness
A central-ui assistant drawer and expanded view, an OpenWebUI chat patch, a top-level sign-in handoff, and an initial message bridge.
Did right
- Kept one persistent iframe so a conversation could survive drawer close and expansion, and kept ERPNext sign-in outside the iframe.
- Added seven focused unit cases for parent-side message source/origin validation and cache-invalidation mapping.
Went wrong
- The child accepted a parent origin claimed in a message and sent its initial handshake to *, while message listeners were not cleaned up.
- The loading/auth state could report success for an error page, the embedded surface was not a verified real chat-only mode, and no live model-to-ERPNext tool flow was demonstrated.
Tests, issues, and UI
- Thinking
- max
- Power
- Not measured during this run.
- Test quality
- Partial. Seven new bridge unit cases cover parent-side source/origin rejection and mutation-to-cache mapping. They encode one invalid query prefix and do not cover child-side origin trust, wildcard handshakes, listener cleanup, authentication return timing, malformed URLs, real embed chrome, model access, or a live ERPNext tool flow.
- Backend and integration issues
- The pending-review lookup limits events to 25 before filtering for OpenWebUI conversations. Newer manual review events can therefore hide an older eligible assistant review. The pure money helper also lacks an ERP-rounding boundary test.
- Verification
- At the fair cutoff, the frontend check and assistant-shell browser test were reported green, but the assistant checks were deterministic or mocked and no live local-model plus ERPNext tool flow had run. Direct inspection confirmed the origin-trust, listener, auth-state, embed, and deployment gaps. All later technically guided fixes and live diagnostics are excluded.
- Human help
- Two neutral completion and verification prompts are included. Three later technical finding rounds, and every resulting change, are excluded from this result.
Qwen 3.8 Flash Next NVFP4
NVFP4 quantization, xhigh reasoning, agent harness
A defensive server-side allocation implementation with extensive backend tests.
Did right
- Passed 60 backend checks and 249 frontend tests.
- Included oldest-first allocation and explicit over-allocation rejection while keeping a clean repository history.
Went wrong
- The chat-side JavaScript bridge was absent, so the assistant UI could not reach the server implementation.
- Its browser suite required unavailable deployment secrets and could not be independently rerun.
Tests, issues, and UI
- Thinking
- xhigh
- Power
- Not measured during this run.
- Test quality
- Detailed Decimal money tests cover invalid inputs, ordering, ties, remainders, non-mutation, and invariants, with additional permission and confirmation contracts. Several integration claims are source-text assertions, and the missing chat bridge left the browser path untested.
- Backend and integration issues
- No backend defect was confirmed in the reviewed allocation path. The blocking defect was integration-level: the assistant UI had no chat bridge to reach the server implementation.
- Verification
- Backend 60/60 and frontend green; browser evidence was not independently runnable; direct inspection confirmed the missing chat bridge.
- Human help
- One completion nudge; no technical steering.
Qwen 3.8 27B BF16
Unquantized BF16 weights, agent harness
A backend slice whose available unit checks passed, plus partial permission and payment-entry work.
Did right
- Passed 40 backend checks.
- Kept permission and payment-entry safeguards in the implementation rather than removing them to make the path easier.
Went wrong
- The frontend failed TypeScript, the assistant patch failed to apply, and the image did not build.
- The session needed twelve extra turns and still ended with a dirty, incomplete working tree before browser verification.
Tests, issues, and UI
- Thinking
- xhigh
- Power
- Not measured during this run.
- Test quality
- The pure allocation tests cover ordering, partial payments, advances, determinism, and basic invariants. They do not cover retry idempotency or the mutation of stored task values, and the frontend and artifact never reached a green gate.
- Backend and integration issues
- Automatic allocation mutates the request values after the idempotency comparison. Retrying the same request ID with the original oldest-first intent can be rejected as different values.
- Verification
- Backend 40/40; frontend typecheck failed; artifact build failed; browser work was not reached; repository remained dirty.
- Human help
- Twelve turns: permission, harness, and resume unblocks, one security-boundary reminder, and one instruction to stop looping and finish.
Gemma 4 31B BF16
Unquantized BF16 weights, default model reasoning, agent harness
A narrow backend slice whose available unit checks passed.
Did right
- Passed 31 backend checks.
- Requested no human assistance.
Went wrong
- The frontend had lint and parser failures, no assistant integration was completed, and no commit was made.
- The write path at the centre of the brief was absent, so silence could not be treated as autonomy.
Tests, issues, and UI
- Thinking
- Default model setting; no explicit effort level was selected.
- Power
- Not measured during this run.
- Test quality
- No feature-specific unit test was added for the new automatic-allocation path. The passing backend count belongs to the pre-existing suite and does not verify Gemma’s new logic.
- Backend and integration issues
- Automatic allocation silently overwrites explicit allocation rows instead of rejecting the ambiguity, and it mutates request values in a way that breaks same-request idempotent retries.
- Verification
- Backend 31/31; frontend failed with ten errors and three warnings; no browser or assistant artifact existed; repository remained dirty.
- Human help
- No extra turns, but the run stopped early.
Scope and limitations
- This is one task, one codebase, one run per model, and one reviewer—not a universal model benchmark.
- Serving configurations and the access incidents encountered by each session differed.
- The comparison does not measure throughput, token cost, or quantization quality.
- A working result for this brief is not proof of production readiness for every workflow.


