← Research Research
Formal Methods & Verified Science
122 tracked items · 25 in the last 30 days · ↓ cooling
Items
-
reddit.com
-
Palomar Opens a Lean Registry That Checks Proofs and Their Descriptions
-
Automated Reasoning: Teach Machines to Think Beyond Prediction
-
AWS's Varun Pant: AI Agents Write Code, Lean4 Must Prove It
-
Research
-
FormalScience: Scalable Human-in-the-Loop Autoformalisation of Science with Agentic Code Generation in Lean
-
arxiv.org
-
Make your project ready for agentic engineering
-
Claude Fable 5.1 made me a really nice animated pelican
-
How Claude Watermarks AI-Generated Text
-
Microconferences
-
The Annals Challenge
-
Palomar – a registry of Lean verified mathematics
-
Bill Gates says we’ve passed AI’s danger thresholds. Now what?
-
Looking for Missed Alarm Bugs in a Formal Verification Tool
-
The Case Against Formal Verification, 50 Years Later
-
Parsing the Stream: A Live Trace Model for Long-Horizon Agents and Their Observers
-
CordisBench: Can Language Models Reason About Component Lifecycles in Dynamic Agent Harnesses?
-
Beyond Scores: Understanding LLM-as-a-Judge Mechanisms in Summarization Evaluation
-
One-shotting a Raccoon Heist game using Claude Fable 5
-
A digestion of the proof of Sendov’s conjecture
-
AI for science needs reasoning, not just data
-
Launch HN: Discovered Materials (YC P26) – AI agents to discover new materials
-
Trump’s AI protectionism has come for robotics
-
Reflections on ICML 2026
-
Sensitivity, Causality, and Repair Dissociate: A Layer-Wise Analysis of Perturbation Robustness and Its Scaling
-
Test-Time Scaling in Reasoning LLMs: Inference Regimes, Evaluation, and Reproducibility
-
Stateless MCP has recaptured my interest (and inspired mcp-explorer and datasette-mcp)
-
Circular financing ain’t what it used to be
-
OpenAI’s amazing — but vastly oversold — new model Astra
-
Public Intelligence
-
Open and Shut
-
Building the enterprise environment for agentic AI
-
From Japan, Products the World Will Use: An Interview with Sakana AI’s Head of Product Development
-
Sakana AI、日本語特化のLLM API「Sakana Namazu」を提供開始
-
Omega-seq: ultra-low-background RNA sequencing with faithful molecular counting and precise transcript-end capture
-
The gate of self-address: where decidable adjudication ends
-
FriendBench: Benchmarking Dyadic Familiarity Inference in Humans and Multimodal Large Language Models
-
Universal BCI Personalization: One API for Frozen EEG Trunks and Foundation Models
-
CausalForge: A Formally Grounded, Self-Improving Agentic Framework for Automated Research in Causal Inference
-
When AI Enters the Architecture Design Loop, What Counts as a Contribution?
-
From reproducible evidence to safer and more useful recommendations: nine new ACM TORS articles
-
ARGOS Geopolitical Risk Engine
-
76 Quantum Theorems Completed By AI In New Lean 4 Benchmarks
-
arxiv.org
-
arxiv.org
-
arxiv.org
-
LeanDojo: AI
-
lean-theorem-proving-guide | Skills Marketplace
-
The Lean Theorem Prover: Design, Evolution, and Impact