Kaelum Bench v0.2 · Public Calibration

Kimi K2.6

Accounting Capstone · Run evidence

Provisional · Hybrid

Run identity

Model and configuration

Model display name
Kimi K2.6
Exact model ID
moonshotai/kimi-k2.6
Provider
ModelRun
Exact provider route
openrouter/modelrun/fp4
Runtime ID
openrouter-tool-free-non-streaming-v8
Runtime version
v8
Reasoning profile
reasoning-max-supported
Reasoning configuration
reasoning-max-supported
Tool access
None
Suite
Accounting Capstone
Validity
Provisional
Material configuration parameters
allow_fallbacks
false
data_collection
deny
max_attempts_per_task
1
max_concurrency
1
max_generation_metadata_lookups
1
max_output_tokens
34000
max_raw_response_bytes
16777216
max_request_bytes
1048576
metadata_timeout_seconds
60
request_timeout_seconds
300
require_parameters
true
temperature
0
zdr
true

Published score

Scoring evidence

Score
918.9 / 1000
Percentage
91.9%
Scoring method
Hybrid
Comparison protocol
Configuration-specific
Tested-model cost
$1.92
Judge cost
$2.90
Total cost
$4.81

Usage

Tokens and execution

Input tokens (tested model)
91,991
Output tokens (tested model)
547,020
Total tokens (tested model)
639,011
Attempted tasks
100 / 100
Valid responses
99 / 100
Invalid JSON
1
Schema failures
0
Provider failures
0
Missing responses
0

Judge evidence

Independent assessment

Judge model
openai/gpt-5.6-sol
Judge provider
OpenRouter (Azure EU)
Judge route
Not publicly released
Judge region
Not publicly released
Judge input tokens
239,925
Judge output tokens
47,852
Judge tokens
287,777
Judge assessments
99
Assessment not required
0
Judge retries
0

Release provenance

Versioned run evidence

Benchmark ID
FMB-CAP-ACC-03
Benchmark version
0.3.0-dev
Scorer ID
fmb-live-scorer-v1
Configuration key
config-a026984a6b07490c6062157850f436aef8641a483028dffd95d14bd915367935
Campaign ID
03700801-ab9d-510f-918b-09631c7a5164
Comparison group ID
accounting-capstone-0.3.0-dev-configuration-specific-v1
Started at
Completed at
Run ID
7edb131b-7965-5478-8ffe-f5b8d29dc54d
Signed release ID
9d07d6d8-80f4-4bca-9adf-1f436336d09d
Release ID
80817257-aa7f-4fb9-867b-66e9ed813de9
Manifest SHA-256
d82a32b09acf027f94de5fe162ca71dabb3a389a4e2b12f58fa3d1faec73a3c8
Source aggregate
fmb-live-runs-capstone.json
Source aggregate SHA-256
922b6a973199cf2ce4c494392dce2e6636b459127a51b6ad4a065190241fb9bb
Source aggregate bytes
17,397

Authenticated publication

Signed release evidence

The sanitized single-run JSON is a derived projection and is not separately signed. The Ed25519 signature authenticates the exact public manifest bytes, including the release inventory. The manifest in turn binds this run's source aggregate by its SHA-256 digest and byte size.

Verify this evidence
  1. Verify the detached Ed25519 signature over the domain-prefixed exact manifest bytes using the published public key.
  2. Confirm that the sidecar payload SHA-256 equals the manifest SHA-256 shown above.
  3. Hash the source aggregate and compare its SHA-256 and byte size with the signed manifest entry.
  4. Resolve this run ID through the public run index and match its campaign ID inside that aggregate.
node -e 'const fs=require("node:fs"),c=require("node:crypto"),m=fs.readFileSync("public-manifest.json"),s=JSON.parse(fs.readFileSync("public-manifest-signature.json","utf8")),k=fs.readFileSync("kaelum-bench-public-release-20260718-v1.pem");console.log(c.verify(null,Buffer.concat([Buffer.from("kaelum-bench-fmb-public-release-v1\0"),m]),k,Buffer.from(s.signature_base64,"base64")))'