Benchmark2025-07-08Evidence A
Perturbation models meet a stronger baseline
A Nature Methods benchmark found evaluated deep-learning perturbation models did not outperform simple linear baselines on the tested datasets.
Why it mattersIt resets the evidence bar from model scale to prospective, cross-context utility.
Virtual-cell products need prospective validation and decision-linked benchmarks, not architecture claims.Infrastructure2025-06-01Evidence C
Arc organizes a shared virtual-cell evaluation layer
Arc Institute describes a Virtual Cell Atlas, perturbation datasets and a public challenge around cell-response prediction.
Why it mattersShared data and evaluation can make competing approaches comparable.
The infrastructure layer may be more durable than any single current model.Research2023-10-19Evidence A
Biomechanics moves from lab equipment to smartphone video
OpenCap reported validation of kinematics and dynamics estimated from two or more smartphone videos.
Why it mattersLower-cost capture expands the environments and populations that can be measured.
The near-term opportunity is measurement infrastructure with explicit task-specific error bounds.Research2023-12-20Evidence A
Closed-loop laboratories demonstrate bounded autonomy
Coscientist and A-Lab demonstrated automated planning or execution in tightly specified chemistry and materials workflows.
Why it mattersThey provide concrete evidence beyond software-only agent demos.
The credible wedge is bounded, instrumented workflows with audit trails—not a general autonomous scientist.