Public Calibration
Accounting Capstone
Kaelum Bench · FMB-CAP-ACC-03
This calibration view covers the concrete 100-task suite only. Its provisional 1000-point hybrid score combines deterministic and rubric components, applied modifiers, and an independent AI judge. It exposes only sanitized score, execution, token, cost, and judge-identity aggregates. It contains no prompts, responses, expected solutions, rubrics, task IDs, provider generation IDs, keys, rationales, digests, or private file paths.
Accounting Capstone · 100-task sub-benchmark
Calibration results
Sanitized lifecycle, coverage, score, token, and cost aggregates for FMB-CAP-ACC-03. Results are provisional and specific to the recorded model configuration.
- Sub-benchmark
- Accounting Capstone · FMB-CAP-ACC-03
- Models
- 1 calibration run
- Coverage
- 100 tasks per model
- Deterministic
- 800 deterministic points
- AI rubric
- 200 AI-judged rubric points
Execution detail
Model runs
Provisional hybrid results combine deterministic scoring, an independent AI judge, and applied modifiers.
Provisional · Configuration-specific calibration · Single runs; no statistical ranking.
OpenRouter (moonshotai/int4) · max reasoning
Kimi K3 — open run evidence
moonshotai/kimi-k3
Judge openai/gpt-5.6-sol · OpenRouter (Azure EU)
100 / 100
95.3%
952.8 / 1000
769.0 / 800 deterministic
192.0 / 200 AI rubric
-8.2 modifier adjustment
100 valid answers
0 invalid JSON responses
100 judge assessments
3 judge retries
493,806 total tokens
$4.33 total cost