Improvement Control Center

Discovery, candidates, integration, probation, and outcomes
Connecting
Back to Hiro
Current operating model
Continuous improvement flow
Standing authority · report afterward
1QueueRank discoveries, friction, failures, and curiosity ideas
2InvestigateActively reproduce the opportunity locally
3ConstructBuild one bounded external-worktree candidate
4TestTargeted, public, held-out, invariant, and latency evidence
5Synthetic soakRepeated isolated probes: low 15m · moderate 60m
6ActivateDurable fast-forward, hidden restart, and loaded-revision verification
7Runtime probationEight hours of live health checks with automatic additive rollback
8FinalizeMark implemented only after verified runtime probation
Live transactions
All active improvement work
Continuous authority
Governor and queue health
0 active
Strategic control layer
Open harness research agenda
Read only · no promotion authority
Problems persist across observations, competing hypotheses, and failed experiments. Progress means changing the evidence state of an important problem—not merely increasing experiment throughput.
Durable problem identities
Prioritized research map
Loading policy
Agenda governance
Review rhythm and authority
Canonical model qualification
Frozen, versioned model leaderboard
Independent from live Hiro
Canonical scores use a sealed reference harness. Live-Hiro compatibility runs are tracked separately and never alter the longitudinal leaderboard.
Registered candidates
Model profiles
Inactive until explicitly run
Qualification identity
Reproducibility boundary
Hash verified
Independent model ledger
All model evidence
Canonical · screen · live integration
Transparent priority
Ranked improvement queue
Showing all
Source leads remain visible in a 100-slot reservoir. Hiro keeps a ten-item external backlog, gives repairable candidates up to three bounded attempts, and never retries safety or invariant failures.
Generational signal
Score and pass-rate trend
ScorePass rate
Latest run
Capability categories
No suite
Audit trail
Recent completed runs
Read only
Recursive improvement
Candidate evidence
Automatic isolated integration
Latest diagnostic
Nightly improvement signal
Public development suite
Longitudinal evidence
Nightly run history
Held-out excluded
Sandbox audit trail
All autonomously tested changes
No automatic promotion
What requires approval?
Raw observations never require approval. Approval applies only to a proposed offline evaluation case with a finding, prompt, assertions, and frozen snapshot hash. Approval never authorizes promotion or Stage 6.
Sanitized public evidence
Internet observation snapshots
Verified · read only
Baseline failures
Pending evaluation case proposals
Human decision required
Explicit decisions
Snapshot-bound evaluation cases
Offline replay only
Router activity
Task, model, time, and token usage
Prompt-free telemetry
Token allocation
Usage by model
Input + output
Experiment ledger
All recent evaluations
Promotion evidence
Candidate versus baseline
Proposal only
Strange-loop ancestry
Variant lineage and fingerprints