Weights that never freeze. A continual-learning architecture, and the pre-registered experiment that falsified its central claim. https://arxiv.org/abs/2607.20792
-
Updated
Jul 22, 2026 - Python
Weights that never freeze. A continual-learning architecture, and the pre-registered experiment that falsified its central claim. https://arxiv.org/abs/2607.20792
A negative result on joint-embedding predictive architectures for prediction markets. The martingale-collapse diagnostic is sound and predicts nothing, more training makes the representation worse, and the metric everything was ranked by was mostly measuring input reconstruction. Pre-registered gates, full findings record.
Policy-Gated Governance for Operational AI Agents
JEPA agent playing Minecraft from pixels: latent world model + MPC planning, 664K params on one 8GB GPU, trained on raw gameplay with no labels. A complete lab notebook - including a 20-attempt research dead end, documented with its root cause.
One fixed 9-stage QSAR and conformal-prediction drug-discovery pipeline run unchanged across ENPP1, NLRP3, TYK2 and IRAK4 — its most valuable outputs are the two refusals.
Distribution-shift teardown of openpilot v0.9.7 supercombo: does a production L2 self-driving model know when it's blind? (it doesn't, and silently)
[zenodo.20574533] CGAA: Concept Guided Adversarial Attacks
A $0 falsification lab across two markets — crypto (~111 hypotheses) and Polymarket prediction markets (172,830 resolved markets, 1.36M trades). 184+ techniques through one committed anti-overfitting gauntlet. 0 survive to a deployable edge; the reusable validation harness is the asset. Agent-ready. MIT.
"Publication bias and the canonization of false facts" published in eLife (2016)
Reproducible evaluation harness for hidden coordination variables in multi-agent LLM systems.
Does cache-aware MoE routing degrade generation, and would the usual metrics notice? 432 pre-registered generations. Every number traced to a result file.
从20个项目中系统提取可复用方法论模式的实验记录。包含10轮正式审查实证(4后端)、58项发现、G5可追溯审计。不成熟框架,诚实的实验记录。
Measurement-first research on local Mixture-of-Experts inference under a hardware contract. 3 measured laws, 4 falsified ideas, and paper site.
Research platform for discovering and rigorously falsifying crypto trading strategies: event-sourced paper trading, implementation-parity verification, and a documented negative-results record.
Does gnomAD LOEUF constraint predict drug-target safety? A negative result: LOEUF measures genetic loss-of-function tolerance well and clinical safety of inhibition poorly.
Component-level benchmark for catastrophic forgetting in world models. Two negative results: forgetting does not follow the labelled task-distance axis, and it happens in the encoder, where the usual metrics cannot see it. 375 runs, with code, data and paper.
Reproducible Apple Silicon benchmark: adaptive 2/4-token prompt lookup did not beat fixed-2 on Qwen3-0.6B.
Signed attention for transformers — sinh/cosh attention giving weights in [-1,1] instead of softmax's positive-only, made FlashAttention/SDPA-compatible by channel doubling. Includes a from-scratch softmax-vs-SBA comparison and a documented negative result.
A global, open-source registry and standardized schema for null findings, non-significant outcomes, and failed trials in medical research. Built to eliminate publication bias and accelerate biomedical discovery.
A negative-result study and a falsification protocol for LLM memory systems.
Add a description, image, and links to the negative-results topic page so that developers can more easily learn about it.
To associate your repository with the negative-results topic, visit your repo's landing page and select "manage topics."