Kaelum Bench v0.2 · Public Calibration
Kimi K3
Accounting Knowledge · Run evidence
Provisional calibration result — not a verified v1 result.
Configuration-specific evidence for a single published run. It is not a statistical ranking and is not an overall score.
Run identity
Model and configuration
- Model display name
- Kimi K3
- Requested model ID
- moonshotai/kimi-k3
- Provider
- Moonshot AI
- Reasoning configuration
- max
- Suite
- Accounting Knowledge
- Validity
- Provisional
Published score
Scoring evidence
- Score
- 550.5 / 600
- Percentage
- 91.7%
- Scoring method
- Hybrid
- Comparison protocol
- Configuration-specific
- Tool access
- None
- Published total cost (model + judge)
- $1.89
Usage
Tokens and execution
- Input tokens (tested model)
- 59,952
- Output tokens (tested model)
- 28,331
- Total tokens (tested model)
- 88,283
- Attempted tasks
- 60 / 60
- Valid responses
- 60 / 60
- Invalid JSON
- 0
- Schema failures
- 0
- Provider failures
- 0
- Missing responses
- 0
Release provenance
Versioned run evidence
- Benchmark ID
- FMB-KNOW-ACC-01
- Benchmark version
- 0.1.0-dev
- Scorer ID
- fmb-live-scorer-v1
- Started at
- Completed at
- Run ID
- 41fd5c4a-ad1d-58a0-a72b-163265c413e9
- Signed release ID
- 0bae13a5-a0c9-44fb-b6df-6bc83a6847d5