Kaelum Bench v0.2 · Public Calibration

DeepSeek V4 Flash 0731

Accounting Capstone · Run evidence

Provisional · Hybrid

Single configuration-specific calibration run. Not a verified v1 result or statistical ranking. The model-judged component is one stored outcome. The current public release has no model-judge repeatability qualification. A fresh judge call is not guaranteed to reproduce it.

Methodology →

Run summary

Published result

Score

78.8/ 100

Raw score 787.5 / 1000

Valid outputs
95 / 100

Successfully scored responses

Model cost
$0.08

Tested-model inference only

Judge cost
$11.93

Independent AI evaluation

Total cost
$12.01

Model plus evaluation

Scoring method
Hybrid

Configuration-specific

Release status
Provisional

Public calibration

Cost composition

Recorded run cost

Model inference · $0.08AI evaluation · $11.93

Run identity

Model setup

Model
DeepSeek V4 Flash 0731
Provider
DeepInfra
Exact model ID
deepseek/deepseek-v4-flash-0731
Benchmark
Accounting Capstone
Reasoning
max
Tool access
None

Output health

Execution quality

Valid outputs

95 / 100

95.0%

Attempted
100
Model tokens
528k
Not valid
5

Scoring process

How this run was evaluated

Hybrid scoring combines deterministic checks with an independent AI rubric assessment. The judge evaluates tasks that require rubric-based review; it does not generate the tested model's answers.

No signed public judge-policy record is available for this release.

Judge model
openai/gpt-5.6-sol
Assessments
95 / 100
Judge tokens
581,369
Judge retries
1
Technical appendixFull technical receiptExact configuration, execution, judge and run-identity fields for stored-evidence audit and replay.Open full technical receipt

Run identity

Exact model configuration

Exact model ID
deepseek/deepseek-v4-flash-0731
Provider
DeepInfra
Exact provider route
openrouter/deepinfra/fp4
Runtime ID
openrouter-tool-free-non-streaming-v8
Reasoning profile
reasoning-max-supported
Reasoning configuration
max
Tool access
None

Material configuration parameters

allow_fallbacks
false
data_collection
deny
max_attempts_per_task
3
max_concurrency
3
max_generation_metadata_lookups
1
max_output_tokens
44000
max_raw_response_bytes
16777216
max_request_bytes
1048576
metadata_timeout_seconds
60
request_timeout_seconds
300
require_parameters
true
temperature
0
zdr
true

Scoring and usage

Exact execution evidence

Score
787.5 / 1000
Percentage
78.8%
Scoring method
Hybrid
Comparison protocol
Configuration-specific
Tested-model cost
$0.0849
Judge cost
$11.93
Total cost
$12.01
Input tokens (tested model)
97,742
Output tokens (tested model)
430,296
Total tokens (tested model)
528,038
Attempted tasks
100 / 100
Valid responses
95 / 100
Invalid JSON
5
Schema failures
0
Provider failures
0
Missing responses
0

Judge evidence

Independent assessment details

Judge model
openai/gpt-5.6-sol
Judge provider
OpenRouter (Azure EU)
Judge tokens
581,369
Judge assessments
95
Judge retries
1

Run provenance

Versioned run identity

These timestamps delimit the recorded wall-clock lifecycle span. They do not measure active model execution time or comparable latency.

Benchmark ID
FMB-CAP-ACC-03
Benchmark version
0.3.0-dev
Scorer ID
fmb-live-scorer-v1
Started at
Completed at
Run ID
6f32c95b-abae-53a0-b9f1-438f8a6ddad0

Authenticated publication

Signed release evidence

Most visitors only need the run JSON. The other files let auditors and developers independently verify the publication chain behind this result.

Start with this file

Recommended · this run

Single-run JSON

A sanitized, machine-readable copy of this exact result: model configuration, score, usage, execution, costs, judge summary and provenance. Best for reviewing, archiving or importing this run. The sanitized single-run JSON is a derived projection and is not separately signed.

Download sanitized run JSON

Advanced downloads

Release verification files

These files work together. A single verification file is not useful on its own.

Release inventory

Public manifest

The signed inventory of this release. It lists every published file together with its SHA-256 digest and byte size.

Download signed manifest

Authenticity proof

Ed25519 signature

The detached signature over the exact manifest bytes. It lets the public detect whether the signed manifest was changed.

Download Ed25519 signature

Verification key

Public key

The public verification key used to check the signature locally. It cannot be used to create a valid Kaelum signature.

Download public key

All public runs

Run index

A machine-readable directory of all 126 public runs, including their run IDs, benchmark versions and configuration references.

Download public run index

Benchmark dataset

Source aggregate

The complete published dataset for Accounting Capstone, covering all model runs in this benchmark. Its hash and byte size are bound by the signed manifest.

Download source aggregate
Cryptographic verification detailsPublication identifiers, hash bindings and a local verification command.Open cryptographic verification details

The Ed25519 signature authenticates the exact public manifest bytes, including the release inventory. The manifest in turn binds this run's source aggregate by its SHA-256 digest and byte size.

Configuration key
config-b2d14de9f9b48ff5b72675844c905dffc89d278c42a8fbf59672f04bdffeb2f5
Campaign ID
74566f95-41ff-521f-b0a5-8fea59044c7b
Comparison group ID
accounting-capstone-0.3.0-dev-configuration-specific-v1
Signed release ID
2459535f-f982-5826-81ed-c8f6d1514773
Release ID
c44af57d-9091-45c8-9299-96dee423e175
Publication key ID
kaelum-bench-public-release-20260718-v1
Manifest SHA-256
caa004fcccade516afe661a318e69b555844595cbcd726cdc749ff3c64a6378f
Source aggregate
fmb-live-runs-capstone.json
Source aggregate SHA-256
85402a8acfbe9f2ff893e95de5c1e105daada9d13ebf267a87fc67d858020660
Source aggregate bytes
21,048

Verify the evidence chain

  1. Verify the detached Ed25519 signature over the domain-prefixed exact manifest bytes using the published public key.
  2. Confirm that the sidecar payload SHA-256 equals the manifest SHA-256 shown above.
  3. Hash the source aggregate and compare its SHA-256 and byte size with the signed manifest entry.
  4. Resolve this run ID through the public run index and match its campaign ID inside that aggregate.
node -e 'const fs=require("node:fs"),c=require("node:crypto"),m=fs.readFileSync("public-manifest.json"),s=JSON.parse(fs.readFileSync("public-manifest-signature.json","utf8")),k=fs.readFileSync("kaelum-bench-public-release-20260718-v1.pem");console.log(c.verify(null,Buffer.concat([Buffer.from("kaelum-bench-fmb-public-release-v1\0"),m]),k,Buffer.from(s.signature_base64,"base64")))'