XScientist protocol for research provenance: Researchers created XScientist, a version-control protocol modeled on Git that tracks scientific processes as content-addressed objects across an exploration graph. The system records intermediate hypotheses, negative outcomes, and verification gates into inspectable packages designed to support continuation across automated research workflows.
Automated materials synthesis with an LLM agent: A laboratory platform combines commercial robotic arms, pipettes, sensors, and 3D-printed fixtures with a language-model agent that generates instrument control code. The system synthesized the metal-organic framework ZIF-8 while recording parameter data and linking pipette dispensing rates to resulting particle size distributions.
Resource allocation frameworks for scientific agents: Researchers designed an economic infrastructure model that applies market principles to multi-agent scientific systems. The framework establishes rules for distributing physical and computational budgets, assigning credit, and managing operational liability when autonomous agents request laboratory resources.
Pipeline for scientific knowledge extraction: A computational framework couples lexical search with dense transformer embeddings using tools such as Neo4j and Elasticsearch to extract, structure, and retrieve information from heterogeneous sources. Testing yielded an overall pipeline F1-score of 0.895 across acquisition and knowledge discovery benchmarks.
Capability assurance in self-driving laboratories: A three-tier architecture adds an operational layer between experimental planners and physical lab instruments. This layer translates experimental protocols into formal machine requirements and tracks physical sensor data, invalidating tasks or initiating recalibration if operating conditions deviate from specifications.
Graph Theory Agent: Researchers paired a frozen language model with external planning components and an input representation selector. This configuration raised model scores on algorithmic graph benchmarks and transferred across multiple separate graph problem sets without parameter updates.
Automated derivation of analytical models: The Rosetta multi-agent system extracts mathematical specifications and executable Python models directly from computer architecture papers. The workflow uses verification checks to validate logic and identify unstated assumptions within published manuscripts.
Additional developments:
- A review cataloged agentic AI frameworks designed for autonomous chemical synthesis planning and execution.
- An infrastructure roadmap outlined operational stages for scaling distributed autonomous discovery networks.
- Researchers categorized generative AI workflows and multi-agent systems designed for geographic information science.
- Robotic-assisted percutaneous coronary intervention procedures demonstrated higher stent placement accuracy alongside reduced radiation exposure for operators.
- Machine learning algorithms automated intravascular imaging interpretations and fractional flow reserve calculations.
Sources
- XScientist: A Git-Like Research Protocol for Long-Running Autonomous Scientific Discovery
Autonomous research systems can generate plausible papers while losing the decisions, failed branches, and evidence needed to inspect or continue the work. We present XScientist, a local-first, git-like protocol that treats research state, rather than a manuscript, as the unit of continuation. Hypotheses, experiment attempts, observations, claims, reviews, and handoffs are represented as typed, content-addressed objects in an exploration graph. Immutable checkpoints, explicit negative outcomes, claim–evidence closure, replay boundaries, and authority-aware gates make each transition…
- Quantitative control and recording of materials-synthesis processes using an automated experimentation platform
Data-driven materials development requires the collection of large amounts of high-quality materials data. Full autonomy of materials experiments is anticipated, but its technical hurdles are high and its adoption remains limited. In this study, we constructed a simple, easy-to-deploy automated experimentation platform that focuses not on full autonomy but on the reliable automation and quantitative recording of experimental processes. Specifically, commercially available instruments such as robot arms, electric pipettes, web cameras, and an electronic balance are combined, components such as…
- Agentic Economies for Autonomous Scientific Discovery
Recent advances in agentic Artificial Intelligence (AI) systems have marked a shift in AI for Science: moving away from the use of individual AI systems for narrow task execution, toward multi-agent systems capable of orchestrating complex, end-to-end research workflows and performing (semi-)autonomous scientific discovery. The development of multi-agent AI-for-science systems has primarily focused on improving the cognitive capabilities of AI systems, specifically by making advanced reasoning and hypothesis generation more reliable. However, focusing only on cognitive capability improvement…
- A Computational Framework for Autonomous Knowledge Acquisition, Representation, Retrieval, and Discovery
This study developed and quantitatively evaluated “A Computational Framework for Autonomous Knowledge Acquisition, Representation, Retrieval, and Discovery” to provide an integrated mechanism for processing heterogeneous information and transforming it into structured, retrievable, and discoverable knowledge. A quantitative computational design-and-evaluation approach was employed in which the framework was implemented using Python and evaluated across autonomous knowledge acquisition, knowledge representation, semantic retrieval, and knowledge discovery. The computational environment…
- From Experiment Execution to Verifiable Capability: An Assurance Architecture for Self-Driving Laboratories
Self-driving laboratories increasingly select experiments, orchestrate resources, execute protocols and update scientific models. Yet existing architectures often under-specify a critical runtime question: what evidence shows that the current physical system can satisfy a particular task under its operating conditions? Device availability, command completion and internal feedback do not establish that delivered volume, end-effector pose or sample temperature meets task-specific limits. Periodic calibration provides essential baseline evidence, but cannot support every operation through a…
- GTA: Graph Theory Agent and Benchmark for Algorithmic Graph Reasoning with LLMs
Large Language Models (LLMs) are increasingly asked to reason over structured data such as graphs, yet how reliably they can carry out multi-step graph algorithms in language remains unclear. Existing evaluations tend to use simple tasks on small graphs, to score code generation rather than reasoning over the graph itself, or to fix a single input format. We introduce Graph Theory Bench (GT Bench), a benchmark covering 24 classical graph problems in 44 task-structure settings, with over 100,000 examples across four representations: natural language, structured language, adjacency list, and…
- Rosetta: Automating First-Principles Performance Modeling Using Multi-Agent LLMs
Analytical performance models --- derivations of throughput or speedup from hardware parameters --- make claims independently verifiable and expose binding constraints, yet rarely accompany architecture papers because building one by hand takes weeks of expert effort. We present Rosetta, a multi-agent LLM pipeline that automatically generates first-principles analytical models from research paper PDFs. Given a paper as sole input, Rosetta produces a mathematical specification, an executable Python model, and a plain-English interpretation --- all autonomously, with zero human intervention.…
Also this week
- Perspective Chapter: The Digital Cath Lab – Robotic-Assisted PCI and AI in Lesion Assessment
- Toward autonomous chemical discovery: agentic AI in organic synthesis and nanochemistry
- From Automated Beamlines to Autonomous Science
- From Large Language Models to Agentic Systems: A Systematic Review of Generative AI in GIScience