Kaelum Bench v0.2 · Public Calibration

Qwen3.6-35B-A3B

Accounting Capstone · Run evidence

Provisional · Hybrid

Provisional calibration result — not a verified v1 result.

Configuration-specific evidence for a single published run. It is not a statistical ranking and is not an overall score.

Run identity

Model and configuration

Model display name
Qwen3.6-35B-A3B
Requested model ID
qwen/qwen3.6-35b-a3b
Provider
AkashML
Reasoning configuration
reasoning-max-supported
Suite
Accounting Capstone
Validity
Provisional

Published score

Scoring evidence

Score
744.0 / 1000
Percentage
74.4%
Scoring method
Hybrid
Comparison protocol
Configuration-specific
Tool access
None
Published total cost (model + judge)
$3.49

Usage

Tokens and execution

Input tokens (tested model)
84,351
Output tokens (tested model)
378,298
Total tokens (tested model)
462,649
Attempted tasks
100 / 100
Valid responses
100 / 100
Invalid JSON
0
Schema failures
0
Provider failures
0
Missing responses
0

Release provenance

Versioned run evidence

Benchmark ID
FMB-CAP-ACC-03
Benchmark version
0.3.0-dev
Scorer ID
fmb-live-scorer-v1
Started at
Completed at
Run ID
7317f280-80ca-5d17-a18f-1df67b79afb2
Signed release ID
9d07d6d8-80f4-4bca-9adf-1f436336d09d