blog · qwen 3.8 max · 2026-09-30

How Qwen 3.8 Max reads evidence nobody wrote a parser for

Kanmani checks whether the payment an ERC-8004 agent's evidence document claims actually happened on the chain it names. That check needs to know which fields hold the chain and the transaction. Qwen finds them in shapes we have never seen. The verifier refuses any answer it cannot confirm.

the problem Qwen solves

ERC-8004 feedback points at an evidence document and says nothing about its shape. Every operator invents one. To test a payment claim, the verifier has to know that settlement.tx_hash is the transaction in one document and payment.proof is it in another. The mapping from a shape to those fields is what we call the atlas.

On Monad the atlas is five hand-written mappings plus a rule for plain-text comments. Together they read all but 3 of the Monad evidence documents we have checked. That number was 0 on 2026-09-28. Then a new agent wrote a settlement-receipt shape nobody had mapped, so the gap is live and it is exactly what the loop below is for. BNB Smart Chain carries the same registries and far more evidence. We sampled 320 of its documents and the atlas, built entirely from Monad, read none of them. Writing a parser per shape by hand does not scale to a registry where anyone can invent a format tomorrow.

what Qwen does: a tool loop, not a prompt

Qwen 3.8 Max gets the raw bytes of one document and a single tool, try_mapping. The tool runs a proposed mapping through our real evidence parser, the same code that produces published verdicts, and returns exactly what it extracted. Qwen proposes the keys that identify the shape and the path to each payment field, along with what it predicts the parser will pull out. It reads the result and revises. It can also say the shape asserts no payment at all, by setting claim to null.

The loop does not trust the model. It accepts a mapping only when what the parser extracted matches what Qwen predicted. A proposal it cannot confirm costs a turn and can never reach a verdict. Running out of turns is recorded as a give-up, never as a success. Each of those three properties has a test that breaks when its guard is removed (verifier/src/atlas/induce.test.ts).

So the division of labour is fixed. Qwen reads and proposes. Deterministic code decides. No model output touches money. Nothing a model says becomes a true or false without the parser and then the chain agreeing.

run one: a field called value that was not money

recorded 2026-09-28, verifier/src/atlas/inductions/gebo-uptime-report.v1.json

The most common unread shape in the BSC sample carries value: 10000 with valueDecimals: 2. A parser written in a hurry reads that as a payment of 10,000 units, finds no matching transaction and marks the claim false. Qwen converged with one tool call and proposed gebo-uptime-report.v1 with no payment claim, reasoning that the figure is 100.00% uptime and that the document has no transaction, payer, recipient, token or network.

We checked that by hand before believing it. Every document of the shape carries tag1: "uptime" and the same value. A key scan across the sample for anything payment-like returned nothing. There were 41 documents of that shape in the sample. Read wrongly, they would have become 41 invented payment claims, each marked false, on a project whose headline is claims that do not resolve. The value here was an error not made.

run two: live, paid, on Monad

2026-09-30, through the Scribe agent, ERC-8004 #10272

The same loop is a hireable agent. Scribe takes an evidence document over HTTP or A2A and returns a confirmed mapping, paid with x402 upto: the caller signs a ceiling and is charged only for the model turns used.

We sent it a document from a shape that dominates BSC's gzip-encoded evidence (36 of 40 we sampled): a feedback entry with score: 100, tag1: "get top 1 rank -->" and a Telegram link in the comment. Qwen returned feedback-rating.v1, no payment claim. It noted unprompted that the tags were being used for self-promotion rather than for rating. A numeric score of 100 is exactly the kind of field a careless mapping turns into an amount.

ceiling signed
10,000 units (0.01 USDC)
charged
4,000 units: a 2,000 base plus one model turn at 2,000
settled on Monad
0x87a7eb4e92d71e03…

Who paid. The payer and the payee are the same wallet, ours. This proves the path works end to end: a model doing tool-using work behind a metered on-chain payment. It says nothing about demand and we do not count it as revenue.

run three: the loop on Monad's own gap

live, queried on load, the candidate sits on /method

The first two runs were on BSC and both found no payment. This one is on Monad and it asserts a real one, which is the harder case and the more useful. A new agent (ERC-8004 #10255) wrote 3 settlement-receipt feedbacks the hand atlas could not read, so noMappingReadTheShape went from 0 to 3. Each one is a tab-settlement-feedback document that says a payment cleared on Monad.

The trap is a field called value. The document carries value: 100 with valueDecimals: 0 and tag1: "tab" at the top, while the money sits under settlement: the hash in settlement.txHash, the amount in settlement.amount, the token in settlement.asset, the recipient in settlement.collection. A careless mapping reads 100 as the amount, finds no matching transfer, then marks a real settlement false. The loop's job is to point at the settlement fields and prove the parser extracts them.

unmapped on Monad now
3 intact documents
converged proposals waiting
0
candidate mapping
tab-settlement-feedback.v1
proposed by
qwen3.8-max

Before and after, stated plainly. noMappingReadTheShape is 3 right now. The proposed mapping round-trips against the real document, which the offline tests prove. It becomes 0 once a reviewer adds the mapping to the atlas in code, a change that would move 3 Monad claims from unknown to a real true or false. the review queue holds it, with the unmapped claims listed in the open. Nothing on that queue has changed a verdict here.

What a keyed run has and has not proven. The loop, its round-trip gate and its ambiguity guard are proven offline by ten tests with no key and no network. The credited run belongs on Alibaba Cloud Model Studio with qwen3.8-max. Until that key is in place the proposer fills this queue with whichever Qwen 3.8 Max host is configured. GMI Cloud serves the same weights on a separate credential, so this page names which credential actually paid rather than implying.

one provider constraint we hit

The loop originally forced a tool call on every turn with tool_choice: "required". Qwen 3.8 Max rejects that in thinking mode:

InternalError.Algo.InvalidParameter: The tool_choice parameter does not
support being set to required or object in thinking mode

Turning thinking off would have fixed the error and lost the point, because the reasoning is what noticed a field named value was not money. The client now sends auto for this model. The loop already tolerates a turn without a tool call and still accepts nothing it has not put through the parser, so the cost is at most one retry.

what Qwen was worth here and what it was not

Worth: coverage of formats nobody on our side has seen, at the moment a document arrives, without a deploy. The two BSC runs found shapes that assert no payment, which in this domain is the valuable catch, because the expensive mistake is inventing a claim and then publishing that it was false. The Monad run is the other case: a shape that does assert a payment, where the win is reading it from the settlement fields instead of the value trap.

Not claimed: three shapes from three runs are not a coverage figure for any chain. None of the three mappings has been added to the shipped atlas, because one confirmed run is a proposal worth reviewing and not yet a standard. Every Monad verdict on this site still comes from the hand-written atlas.

run it yourself

# the loop directly, with your own key (GMI Cloud or DashScope, OpenAI-compatible)
export GMI_API_KEY=...
export QWEN_BASE_URL=https://api.gmi-serving.com/v1
export QWEN_MODEL=Qwen/Qwen3.8-Max
npx tsx verifier/src/atlas/demo.ts

# or hire Scribe: an unpaid call answers 402 with its x402 options
curl -s -X POST https://kanmani.xyz/api/service/scribe \
  -H 'content-type: application/json' \
  -d '{"document":{"value":"10000","valueDecimals":2,"tag1":"uptime"}}'

Scribe's registration, A2A card and price are on /hire. The loop, the parser and both run records are in the repository under verifier/src/atlas/.