Kaelum Bench v0.2 · Public Calibration
kimi-k3
Financial Analysis · Run evidence
Provisional calibration result — not a verified v1 result.
Configuration-specific evidence for a single published run. It is not a statistical ranking and is not an overall score.
Run identity
Model and configuration
- Model display name
- kimi-k3
- Requested model ID
- moonshotai/kimi-k3
- Provider
- Moonshot AI
- Reasoning configuration
- max
- Suite
- Financial Analysis
- Validity
- Provisional
Published score
Scoring evidence
- Score
- 399.0 / 600
- Percentage
- 66.5%
- Scoring method
- Deterministic
- Comparison protocol
- Configuration-specific
- Tool access
- None
- Published total cost
- $0.516
Usage
Tokens and execution
- Input tokens (tested model)
- 83,254
- Output tokens (tested model)
- 20,650
- Total tokens (tested model)
- 103,904
- Attempted tasks
- 60 / 60
- Valid responses
- 60 / 60
- Invalid JSON
- 0
- Schema failures
- 0
- Provider failures
- 0
- Missing responses
- 0
Release provenance
Versioned run evidence
- Benchmark ID
- FMB-ANAL-FSA-02
- Benchmark version
- 0.2.0-dev
- Scorer ID
- fmb-live-scorer-v1
- Started at
- Completed at
- Run ID
- 51874bd4-7ece-5434-bb99-dd6f81e32d54
- Signed release ID
- 6f584818-802d-457a-8d7e-5e6106c40d2a