ArcticSwarm Coordinates Long-Horizon Multi-Agent Research: Researchers developed ArcticSwarm, an architecture designed to manage multi-agent coordination during extended research tasks. To prevent agents from settling prematurely on early hypotheses, the framework separates evidence collection from integration via a shared bulletin board, applying gated isolation and review gates across three commitment boundaries. ArcticSwarm scored 82.6% accuracy on BrowseComp-Plus using Qwen 3.5-27B and 73.6% on live-web BrowseComp with GPT-5.
TruthInsightBench Tests Autonomous Scientific Discovery: The TruthInsightBench benchmark evaluates the capacity of autonomous agents to extract scientific findings from data without guided analysis steps. Comprising 40 blind tasks across 10 scientific disciplines, the framework uses an automated model judge to grade 29 artifact-grounded criteria across six dimensions. Four tested coding agents scored between 58.4 and 60.3 out of 100, showing functional code execution alongside deficits in experimental controls and robustness checks.
LLM Workflow Replicates and Extends Economics Literature: An open-source workflow applies large language models to process published replication archives from economics journals. Across 4,452 replication packages from five publications, the system flagged discrepancies in 3,460 articles or appendices, lowered computation times by more than a factor of ten in 496 papers, and generated analytical extensions for 923 papers.
LabAgent Standardizes Method Continuity in Research Labs: Researchers designed LabAgent, a system intended to maintain workflow inheritance and preserve procedural knowledge across laboratory staff transitions. The framework verifies executable skills and logs corrective troubleshooting steps to prevent recurring operational errors. In tests covering drug property prediction and statistical genetics, LabAgent reproduced a published figure and surpassed the accuracy of commercial generalist agents.
Cyber-Physical Biomimicry Framework for Marine Energy: Researchers introduced a framework combining biomimicry designs with multi-agent reinforcement learning to harvest marine energy.
GUIAuditor for Mobile Forensics: Researchers built GUIAuditor, an agentic auditing tool using multimodal language models for post-hoc forensic evaluation of mobile graphical user interfaces.
Sources
- ArcticSwarm: Deferring Early Consensus in Long-Horizon Multi-Agent Research
Multi-agent systems have shown strong performance in domains with reliable verifiers such as coding, where multi-parallel candidate generation selected by a verifier is effective. However, such pipelines would not generalize to open-ended, long-horizon research tasks without a verifier. While majority voting or self-consistency is often used to reach consensus as a proxy verifier, parallel agents repeatedly explore the same evidence, while access to peers' partial findings cause search to converge on an early candidate before alternatives are tested. We present ArcticSwarm, a multi-agent…
- TruthInsightBench: An Evidence-Grounded Benchmark for Automated Evaluation of Open-Ended Scientific Discovery Agents
Autonomous coding agents are increasingly proposed as AI-scientist systems that conduct analyses and write research reports, but executing a prescribed analysis is not the same as making a discovery. Existing benchmarks are configured for reproduction: tasks, data, and rubrics are built around a hidden target study, and recovery of its result is rewarded. We present TruthInsightBench, a benchmark configured for discovery. Its 40 blind tasks, drawn from 40 peer-reviewed studies across 10 scientific domains, expose only a neutral scientific objective and frozen data; source conclusions,…
- An LLM Workflow That Reproduces, Improves, and Extends Published Economics Research
We introduce an open-source workflow that enables an LLM to reproduce, improve, and extend an economics article using the article's published replication package.First, the workflow attempts to reproduce the original calculations, checks for discrepancies with published findings, and performs automated sensitivity analysis.Across 4,452 published replication packages for five economics journals, the workflow flags discrepancies in 3,460 articles or their appendices.Second, the workflow improves the original calculations by using a different implementation or algorithm.In 496 articles, the…
- LabAgent: Customize Any Research Hubs for Scientific Discoveries Using AI Agents
Scientific research is a continuous process that emphasizes inheritance. Methods developed by predecessors are often expanded upon by new researchers to explore more novel and in-depth scientific questions. However, the change of lab staff, such as student graduation, leads to a lack of personnel capable of replicating methods. Methods that have been developed with significant effort and resources cannot be continued. To address these limitations, we propose LabAgent, a reproduce and discovery harness tailored for a lab's continuous work. LabAgent employs two mechanisms to guarantee that all…