Kaelum Bench v0.2 · Public Calibration

Claude Opus 5

Accounting Knowledge · Run evidence

Provisional · Hybrid

Run identity

Model and configuration

Model display name
Claude Opus 5
Exact model ID
anthropic/claude-opus-5
Provider
Anthropic
Exact provider route
openrouter/anthropic
Runtime ID
openrouter-tool-free-non-streaming-v8
Runtime version
v8
Reasoning profile
reasoning-max-supported
Reasoning configuration
reasoning-max-supported
Tool access
None
Suite
Accounting Knowledge
Validity
Provisional
Material configuration parameters
allow_fallbacks
false
data_collection
deny
max_attempts_per_task
1
max_concurrency
1
max_generation_metadata_lookups
1
max_output_tokens
34000
max_raw_response_bytes
16777216
max_request_bytes
1048576
metadata_timeout_seconds
60
request_timeout_seconds
300
require_parameters
true
temperature
0
zdr
false

Published score

Scoring evidence

Score
554.1 / 600
Percentage
92.4%
Scoring method
Hybrid
Comparison protocol
Configuration-specific
Tested-model cost
$0.9834
Judge cost
$1.29
Total cost
$2.28

Usage

Tokens and execution

Input tokens (tested model)
93,274
Output tokens (tested model)
20,683
Total tokens (tested model)
113,957
Attempted tasks
60 / 60
Valid responses
60 / 60
Invalid JSON
0
Schema failures
0
Provider failures
0
Missing responses
0

Judge evidence

Independent assessment

Judge model
openai/gpt-5.6-sol
Judge provider
OpenRouter (Azure EU)
Judge route
Not publicly released
Judge region
Not publicly released
Judge input tokens
115,549
Judge output tokens
19,938
Judge tokens
135,487
Judge assessments
60
Assessment not required
0
Judge retries
0

Release provenance

Versioned run evidence

Benchmark ID
FMB-KNOW-ACC-01
Benchmark version
0.1.0-dev
Scorer ID
fmb-live-scorer-v1
Configuration key
config-ef66eed2187cc09eaac374abe4b2b30dea66e50ad713493b2ce3b90806157fd0
Campaign ID
724a3998-0e06-5bfd-a566-21873d55bb99
Comparison group ID
accounting-knowledge-0.1.0-dev-configuration-specific-v1
Started at
Completed at
Run ID
4141bff2-8b92-507e-954d-9c173bcaa047
Signed release ID
0bae13a5-a0c9-44fb-b6df-6bc83a6847d5
Release ID
80817257-aa7f-4fb9-867b-66e9ed813de9
Manifest SHA-256
d82a32b09acf027f94de5fe162ca71dabb3a389a4e2b12f58fa3d1faec73a3c8
Source aggregate
fmb-live-runs.json
Source aggregate SHA-256
0ef185b35b7d9f9b891a40de2bc9a9cee0febe77a41ee700458b6b92b4bb2cd3
Source aggregate bytes
17,057

Authenticated publication

Signed release evidence

The sanitized single-run JSON is a derived projection and is not separately signed. The Ed25519 signature authenticates the exact public manifest bytes, including the release inventory. The manifest in turn binds this run's source aggregate by its SHA-256 digest and byte size.

Verify this evidence
  1. Verify the detached Ed25519 signature over the domain-prefixed exact manifest bytes using the published public key.
  2. Confirm that the sidecar payload SHA-256 equals the manifest SHA-256 shown above.
  3. Hash the source aggregate and compare its SHA-256 and byte size with the signed manifest entry.
  4. Resolve this run ID through the public run index and match its campaign ID inside that aggregate.
node -e 'const fs=require("node:fs"),c=require("node:crypto"),m=fs.readFileSync("public-manifest.json"),s=JSON.parse(fs.readFileSync("public-manifest-signature.json","utf8")),k=fs.readFileSync("kaelum-bench-public-release-20260718-v1.pem");console.log(c.verify(null,Buffer.concat([Buffer.from("kaelum-bench-fmb-public-release-v1\0"),m]),k,Buffer.from(s.signature_base64,"base64")))'