Kaelum Bench v0.2 · Public Calibration
Qwen3.6-35B-A3B
Accounting Capstone · Run evidence
Provisional calibration result — not a verified v1 result.
Configuration-specific evidence for a single published run. It is not a statistical ranking and is not an overall score.
Run identity
Model and configuration
- Model display name
- Qwen3.6-35B-A3B
- Requested model ID
- qwen/qwen3.6-35b-a3b
- Provider
- AkashML
- Reasoning configuration
- reasoning-max-supported
- Suite
- Accounting Capstone
- Validity
- Provisional
Published score
Scoring evidence
- Score
- 744.0 / 1000
- Percentage
- 74.4%
- Scoring method
- Hybrid
- Comparison protocol
- Configuration-specific
- Tool access
- None
- Published total cost (model + judge)
- $3.49
Usage
Tokens and execution
- Input tokens (tested model)
- 84,351
- Output tokens (tested model)
- 378,298
- Total tokens (tested model)
- 462,649
- Attempted tasks
- 100 / 100
- Valid responses
- 100 / 100
- Invalid JSON
- 0
- Schema failures
- 0
- Provider failures
- 0
- Missing responses
- 0
Release provenance
Versioned run evidence
- Benchmark ID
- FMB-CAP-ACC-03
- Benchmark version
- 0.3.0-dev
- Scorer ID
- fmb-live-scorer-v1
- Started at
- Completed at
- Run ID
- 7317f280-80ca-5d17-a18f-1df67b79afb2
- Signed release ID
- 9d07d6d8-80f4-4bca-9adf-1f436336d09d