Agents in Research
« Agents in Research

#27 — ArcticSwarm research agents, TruthInsightBench, economics replication

September 30, 2026

Sources

  1. ArcticSwarm: Deferring Early Consensus in Long-Horizon Multi-Agent Research
    Multi-agent systems have shown strong performance in domains with reliable verifiers such as coding, where multi-parallel candidate generation selected by a verifier is effective. However, such pipelines would not generalize to open-ended, long-horizon research tasks without a verifier. While majority voting or self-consistency is often used to reach consensus as a proxy verifier, parallel agents repeatedly explore the same evidence, while access to peers' partial findings cause search to converge on an early candidate before alternatives are tested. We present ArcticSwarm, a multi-agent…
  2. TruthInsightBench: An Evidence-Grounded Benchmark for Automated Evaluation of Open-Ended Scientific Discovery Agents
    Autonomous coding agents are increasingly proposed as AI-scientist systems that conduct analyses and write research reports, but executing a prescribed analysis is not the same as making a discovery. Existing benchmarks are configured for reproduction: tasks, data, and rubrics are built around a hidden target study, and recovery of its result is rewarded. We present TruthInsightBench, a benchmark configured for discovery. Its 40 blind tasks, drawn from 40 peer-reviewed studies across 10 scientific domains, expose only a neutral scientific objective and frozen data; source conclusions,…
  3. An LLM Workflow That Reproduces, Improves, and Extends Published Economics Research
    We introduce an open-source workflow that enables an LLM to reproduce, improve, and extend an economics article using the article's published replication package.First, the workflow attempts to reproduce the original calculations, checks for discrepancies with published findings, and performs automated sensitivity analysis.Across 4,452 published replication packages for five economics journals, the workflow flags discrepancies in 3,460 articles or their appendices.Second, the workflow improves the original calculations by using a different implementation or algorithm.In 496 articles, the…
  4. LabAgent: Customize Any Research Hubs for Scientific Discoveries Using AI Agents
    Scientific research is a continuous process that emphasizes inheritance. Methods developed by predecessors are often expanded upon by new researchers to explore more novel and in-depth scientific questions. However, the change of lab staff, such as student graduation, leads to a lack of personnel capable of replicating methods. Methods that have been developed with significant effort and resources cannot be continued. To address these limitations, we propose LabAgent, a reproduce and discovery harness tailored for a lab's continuous work. LabAgent employs two mechanisms to guarantee that all…

Also this week

Full transcript
How can autonomous research swarms coordinate long investigations without settling on conclusions too early? That coordination problem is what we examine today on Agents in Research, covering computational methods across scientific inquiry. Here is how the architecture works. In open-ended scientific workflows, standard multi-agent setups tend to stall at premature consensus. Without deterministic test suites to verify intermediate steps, agents that inspect each other's outputs end up duplicating search directions or converging before testing alternatives. Which explains why ArcticSwarm isolates them. Instead of direct subagent communication, the architecture routes observations to a shared bulletin board behind gated isolation. And integration only takes place at three discrete commitment boundaries. By restricting when agents merge findings, alternative hypotheses remain active through the exploration phase. That boundary structure showed measurable differences on benchmarks. Paired with Qwen 3.5-27B on BrowseComp-Plus, ArcticSwarm posted an 82.6 percent accuracy rate. On the live-web BrowseComp setup using GPT-5, performance reached 73.6 percent. Information retrieval on the web is one hurdle, but testing whether an agent can independently formulate scientific claims directly from data is another. That is where script execution gets separated from discovery. TruthInsightBench isolates that distinction across 40 blind tasks in 10 scientific domains. The agents receive frozen datasets and neutral objectives, with no target conclusions, analysis protocols, or expected numbers provided. An automated model judge then grades 29 artifact-grounded items across six dimensions. Tested across four coding agents, the final scores clustered tightly between 58.4 and 60.3 out of 100. They could write and run code, but failed systematically on experimental controls, robustness checks, and scientific judgment. That gap between code execution and research verification also appeared when an automated workflow audited 4,452 replication packages across five economics journals. The system flagged computational or reporting discrepancies in 3,460 of those articles or their appendices. It also refactored the underlying code, reducing execution times by more than a factor of ten across 496 papers, and generated analytical extensions for 923 studies. That type of audit targets past publications, but laboratory operations face a parallel preservation problem in real time: personnel turnover breaking experimental continuity. LabAgent targets that operational gap by pairing method inheritance with verification protocols, keeping analytical skills runnable over time while logging corrective steps to stop recurring procedural errors. In trials on statistical genetics and drug property prediction, it reproduced published figures and exceeded the task metrics of commercial generalist agents. In separate developments, researchers applied multi-agent reinforcement learning to an AI-driven biomimicry cyber-physical framework for marine energy harvesting. And GUIAuditor applied multimodal language models to post-hoc forensic auditing of mobile graphical user interfaces. We will return next week with more developments in research automation. Until next time, on Agents in Research.