Agents in Research
« Agents in Research

#28 — XScientist research provenance, automated materials platform

October 7, 2026

Sources

  1. XScientist: A Git-Like Research Protocol for Long-Running Autonomous Scientific Discovery
    Autonomous research systems can generate plausible papers while losing the decisions, failed branches, and evidence needed to inspect or continue the work. We present XScientist, a local-first, git-like protocol that treats research state, rather than a manuscript, as the unit of continuation. Hypotheses, experiment attempts, observations, claims, reviews, and handoffs are represented as typed, content-addressed objects in an exploration graph. Immutable checkpoints, explicit negative outcomes, claim–evidence closure, replay boundaries, and authority-aware gates make each transition…
  2. Quantitative control and recording of materials-synthesis processes using an automated experimentation platform
    Data-driven materials development requires the collection of large amounts of high-quality materials data. Full autonomy of materials experiments is anticipated, but its technical hurdles are high and its adoption remains limited. In this study, we constructed a simple, easy-to-deploy automated experimentation platform that focuses not on full autonomy but on the reliable automation and quantitative recording of experimental processes. Specifically, commercially available instruments such as robot arms, electric pipettes, web cameras, and an electronic balance are combined, components such as…
  3. Agentic Economies for Autonomous Scientific Discovery
    Recent advances in agentic Artificial Intelligence (AI) systems have marked a shift in AI for Science: moving away from the use of individual AI systems for narrow task execution, toward multi-agent systems capable of orchestrating complex, end-to-end research workflows and performing (semi-)autonomous scientific discovery. The development of multi-agent AI-for-science systems has primarily focused on improving the cognitive capabilities of AI systems, specifically by making advanced reasoning and hypothesis generation more reliable. However, focusing only on cognitive capability improvement…
  4. A Computational Framework for Autonomous Knowledge Acquisition, Representation, Retrieval, and Discovery
    This study developed and quantitatively evaluated “A Computational Framework for Autonomous Knowledge Acquisition, Representation, Retrieval, and Discovery” to provide an integrated mechanism for processing heterogeneous information and transforming it into structured, retrievable, and discoverable knowledge. A quantitative computational design-and-evaluation approach was employed in which the framework was implemented using Python and evaluated across autonomous knowledge acquisition, knowledge representation, semantic retrieval, and knowledge discovery. The computational environment…
  5. From Experiment Execution to Verifiable Capability: An Assurance Architecture for Self-Driving Laboratories
    Self-driving laboratories increasingly select experiments, orchestrate resources, execute protocols and update scientific models. Yet existing architectures often under-specify a critical runtime question: what evidence shows that the current physical system can satisfy a particular task under its operating conditions? Device availability, command completion and internal feedback do not establish that delivered volume, end-effector pose or sample temperature meets task-specific limits. Periodic calibration provides essential baseline evidence, but cannot support every operation through a…
  6. GTA: Graph Theory Agent and Benchmark for Algorithmic Graph Reasoning with LLMs
    Large Language Models (LLMs) are increasingly asked to reason over structured data such as graphs, yet how reliably they can carry out multi-step graph algorithms in language remains unclear. Existing evaluations tend to use simple tasks on small graphs, to score code generation rather than reasoning over the graph itself, or to fix a single input format. We introduce Graph Theory Bench (GT Bench), a benchmark covering 24 classical graph problems in 44 task-structure settings, with over 100,000 examples across four representations: natural language, structured language, adjacency list, and…
  7. Rosetta: Automating First-Principles Performance Modeling Using Multi-Agent LLMs
    Analytical performance models --- derivations of throughput or speedup from hardware parameters --- make claims independently verifiable and expose binding constraints, yet rarely accompany architecture papers because building one by hand takes weeks of expert effort. We present Rosetta, a multi-agent LLM pipeline that automatically generates first-principles analytical models from research paper PDFs. Given a paper as sole input, Rosetta produces a mathematical specification, an executable Python model, and a plain-English interpretation --- all autonomously, with zero human intervention.…

Also this week

Full transcript
A published paper rarely documents every failed branch an automated system explored to reach its conclusion. Preserving that lineage is our focus today on Agents in Research, where we track artificial intelligence in scientific discovery. Here is the latest work. On Agents in Research, we look first at the gap between generating a readable scientific manuscript and recording the actual steps that produced it. Because an autonomous system can write a complete paper while omitting the rejected branches, the intermediate data, and the specific decision points that an outside researcher needs to reproduce the result. That is the premise behind XScientist. It operates as a local-first, git-like protocol that treats the underlying research state as the primary unit of continuation rather than the final text. Meaning hypotheses, observations, experimental attempts, and review gates are structured as typed, content-addressed objects inside an exploration graph. And it introduces immutable checkpoints, explicit negative outcome tracking, and replay boundaries. Crucially, passing an integrity check within the graph is not treated as equivalent to scientific validity—it simply guarantees external inspectability. The reference implementation packages that entire cycle—planning, execution, review, and repair—into an Agent-Native Research Artifact that another laboratory can verify or resume. Which brings the problem directly to physical execution. Moving from virtual workflows to wet labs introduces integration barriers, but one team built an automated materials experimentation platform using off-the-shelf components. They paired standard commercial hardware—robotic arms, balances, webcams, and electric pipettes—with custom three-dimensional printed fixtures, all orchestrated by a language model agent. The model writes the control code for each instrument and logs the procedure step by step. They tested it on the chemical synthesis of the metal-organic framework ZIF-8. And controlling the pipette dispensing speed systematically turned out to yield reproducible particle size distributions. But execution reliability remains an issue when hardware drifts or misbehaves. Conventional automation only checks whether a machine finished a command or passed its routine calibration. It does not track whether the delivered volume, thermal state, or tool position matched the tolerance required for that specific task. Which is why researchers designed a three-layer capability assurance architecture for self-driving laboratories. It slots directly between high-level task planning and physical hardware. That middle layer translates protocol constraints into machine-readable requirements and tracks them using uncertainty-aware sensors. If a collision or thermal shift occurs, the capability record is invalidated immediately. Forcing the system to pause, reassign, or run verification before anything resumes. Of course, physical constraints are not the only limit—there are also computational and financial boundaries. That friction led to an economic infrastructure framework for allocating resources between artificial intelligence agents and human researchers. When multi-agent systems direct research pipelines, instrument time and budgets become hard bottlenecks. The model uses market mechanisms and institutional rules to set research priorities, assign operational liability when an experiment fails, and handle credit attribution across participating agents. To feed these experimental pipelines, systems also need to process existing literature autonomously. One framework evaluated a pipeline for knowledge acquisition, semantic retrieval, and discovery across unstructured documents. It combined dense transformer embeddings with lexical search algorithms, utilizing tools like spaCy, PyTorch, Hugging Face Transformers, Neo4j, and Elasticsearch. In benchmark tests, it reached an acquisition F1-score of 0.906, a retrieval Precision at 10 of 0.863, and an overall pipeline F1-score of 0.895. At the same time, processing theoretical concepts requires precise computational reasoning. Language models often stumble over graph algorithms depending on how the graph is represented. The Graph Theory Agent, or GTA, separates representation selection and planning scaffolding from a frozen language model executor to fix that inconsistency. Tested on Graph Theory Bench across twenty-four classical problems, it raised Phi-4's accuracy on the easy split from 53.5 percent to 69.1 percent, and on the hard split from 33.0 percent to 41.5 percent. And that scaffolding held its accuracy when transferred to external datasets like GraCoRe and NLGraph without retraining. A related extraction effort is Rosetta, which targets analytical performance models inside computer architecture papers. Deriving first-principles equations for hardware throughput from text normally takes weeks of manual work. Rosetta reads the paper's PDF and produces formal mathematical specifications, executable Python simulations, and structured interpretations. It uses critic agents, best-of-N ensembling, and execution checks. In several cases, those validation checks exposed implicit assumptions and led authors to revise claims in pending manuscripts. Rounding out recent developments: robotic-assisted percutaneous coronary intervention improved stent placement accuracy while lowering operator radiation exposure, and machine learning models automated intravascular imaging analysis and fractional flow reserve calculations. We also saw a review of agentic frameworks for chemical discovery, an infrastructure progression model for distributed autonomous science ecosystems, and an overview of generative and agentic architectures applied to geographic information science. We will return next week with more developments in automated laboratory workflows. Until next time, on Agents in Research.