Kaelum Bench v0.2 · Public Calibration

Kimi K3

Accounting Knowledge · Run evidence

Provisional · Hybrid

Provisional calibration result — not a verified v1 result.

Configuration-specific evidence for a single published run. It is not a statistical ranking and is not an overall score.

Run identity

Model and configuration

Model display name
Kimi K3
Requested model ID
moonshotai/kimi-k3
Provider
Moonshot AI
Reasoning configuration
max
Suite
Accounting Knowledge
Validity
Provisional

Published score

Scoring evidence

Score
550.5 / 600
Percentage
91.7%
Scoring method
Hybrid
Comparison protocol
Configuration-specific
Tool access
None
Published total cost (model + judge)
$1.89

Usage

Tokens and execution

Input tokens (tested model)
59,952
Output tokens (tested model)
28,331
Total tokens (tested model)
88,283
Attempted tasks
60 / 60
Valid responses
60 / 60
Invalid JSON
0
Schema failures
0
Provider failures
0
Missing responses
0

Release provenance

Versioned run evidence

Benchmark ID
FMB-KNOW-ACC-01
Benchmark version
0.1.0-dev
Scorer ID
fmb-live-scorer-v1
Started at
Completed at
Run ID
41fd5c4a-ad1d-58a0-a72b-163265c413e9
Signed release ID
0bae13a5-a0c9-44fb-b6df-6bc83a6847d5