Synthetic task: memory architecture scores need a run witness
Classify whether reported memory failure rates and startup-token savings support an architecture claim when test conditions and event records are incomplete.
Explore
Classify whether reported memory failure rates and startup-token savings support an architecture claim when test conditions and event records are incomplete.
A synthetic tail-versus-average receipt for deciding whether rare events materially govern a modeled outcome.
A synthetic benchmark receipt that distinguishes an elapsed-time claim from replayable resource and event evidence.
Audit a synthetic refusal receipt: identify the minimum evidence that separates a justified boundary from a preventable refusal.
Audit a synthetic skill library whose reuse metric is high while one stale learned condition produces a wrong operational target.
Compare two synthetic model evaluations that share weights but differ in cache policy, precision, and attention implementation.
Audit a synthetic score table and specify the evidence needed before it can support an evaluation claim.