Kaelum Bench v0.2 · Public Calibration

kimi-k3

Financial Analysis · Run evidence

Provisional · Deterministic

Provisional calibration result — not a verified v1 result.

Configuration-specific evidence for a single published run. It is not a statistical ranking and is not an overall score.

Run identity

Model and configuration

Model display name
kimi-k3
Requested model ID
moonshotai/kimi-k3
Provider
Moonshot AI
Reasoning configuration
max
Suite
Financial Analysis
Validity
Provisional

Published score

Scoring evidence

Score
399.0 / 600
Percentage
66.5%
Scoring method
Deterministic
Comparison protocol
Configuration-specific
Tool access
None
Published total cost
$0.516

Usage

Tokens and execution

Input tokens (tested model)
83,254
Output tokens (tested model)
20,650
Total tokens (tested model)
103,904
Attempted tasks
60 / 60
Valid responses
60 / 60
Invalid JSON
0
Schema failures
0
Provider failures
0
Missing responses
0

Release provenance

Versioned run evidence

Benchmark ID
FMB-ANAL-FSA-02
Benchmark version
0.2.0-dev
Scorer ID
fmb-live-scorer-v1
Started at
Completed at
Run ID
51874bd4-7ece-5434-bb99-dd6f81e32d54
Signed release ID
6f584818-802d-457a-8d7e-5e6106c40d2a