Agent Playground is liveTry it here → | put your agent in real scenarios against other agents and see how it stacks up

Key Takeaway

A specialized AI multi-agent system can plan and run complex, cross-modal neuroscience analyses, cutting routine work and producing valid discoveries—it beat general coding agents by up to 17.7% on a domain benchmark.

Core Insights

SeekBrain builds a growing library of field-tested analysis recipes pulled from papers and open code, then uses a two-part agent setup to turn high-level research questions into validated analysis pipelines. The system produced more scientifically valid interpretations than general-purpose agents and is designed for researcher-in-the-loop control so experts can guide and refine outputs. On a 32-task neuroscience benchmark and two real datasets (zebrafish and mouse), it both improved benchmark scores and generated interpretable discoveries such as low-dimensional neural representations and a shared axis of regional decoding strength. Multi-Agent Scientific Research.
Explore evaluation patternsSee how to apply these findings
Learn More

Data Highlights

1Outperformed Claude Code by 11.7% and Codex by 17.7% on the BrainArena 32-task benchmark.
2Removing the Neuroscience Analysis Repertoire caused a performance drop of over 10 points, showing the repertoire is essential for scientific validity.
3BrainArena benchmark: 32 expert-annotated tasks scored with rubrics containing over 700 deduction subitems for figure fidelity, method correctness, and interpretive validity.

What This Means

Neuroscience labs and data teams who want to automate routine analyses and scale expertise across multimodal datasets will gain the most. AI engineers and product leads building domain-specific agent systems can use SeekBrain as a model for embedding curated workflows and validation into agent pipelines. Specialist Agent

Key Figures

Figure 1 : Overview of SeekBrain. a , Overall workflow. SeekBrain processes research query along with neuroscience data through two core engines. The Research Planning Engine generates a tree-structured hierarchical research plan, while the Analysis Execution Engine performs the planned analysis tasks and produces figures and statistical evidence, which are synthesized into a report. Methodological rigor is enforced by the Neuroscience Analysis Repertoire, which contains a wide range of domain-specific analysis recipes crystallized from published studies and their associated codebase. b , Research Planning Engine. After receiving a broad research query, the engine conducts an iterative debate between the Orchestrator and the Reviewer Agents to decompose the query into smaller tasks. This process yields a tree-structured research plan where terminal nodes represent specific analysis tasks for downstream execution. c , Analysis Execution Engine. A Formulation Agent converts each analysis query into a structured computational schema, which the Code Agent then translates into Python scripts. A Validation Agent evaluates the outputs across four core dimensions. This closed-loop refinement process continues until the results pass validation. d , Taxonomy of the Neuroscience Analysis Repertoire, which includes both unimodal and multimodal analysis recipes. Representative recipe types are presented. FC, functional connectivity; SC, structural connectivity. e , SeekBrain’s researcher-in-the-loop paradigm, which enables users to flexibly guide the analysis process and supports the continuous refinement of analysis recipes based on expert interaction traces.
Fig 1: Figure 1 : Overview of SeekBrain. a , Overall workflow. SeekBrain processes research query along with neuroscience data through two core engines. The Research Planning Engine generates a tree-structured hierarchical research plan, while the Analysis Execution Engine performs the planned analysis tasks and produces figures and statistical evidence, which are synthesized into a report. Methodological rigor is enforced by the Neuroscience Analysis Repertoire, which contains a wide range of domain-specific analysis recipes crystallized from published studies and their associated codebase. b , Research Planning Engine. After receiving a broad research query, the engine conducts an iterative debate between the Orchestrator and the Reviewer Agents to decompose the query into smaller tasks. This process yields a tree-structured research plan where terminal nodes represent specific analysis tasks for downstream execution. c , Analysis Execution Engine. A Formulation Agent converts each analysis query into a structured computational schema, which the Code Agent then translates into Python scripts. A Validation Agent evaluates the outputs across four core dimensions. This closed-loop refinement process continues until the results pass validation. d , Taxonomy of the Neuroscience Analysis Repertoire, which includes both unimodal and multimodal analysis recipes. Representative recipe types are presented. FC, functional connectivity; SC, structural connectivity. e , SeekBrain’s researcher-in-the-loop paradigm, which enables users to flexibly guide the analysis process and supports the continuous refinement of analysis recipes based on expert interaction traces.
Figure 2 : BrainArena evaluation. a, Overview of BrainArena construction and evaluation process. Expert-curated tasks are generated from neuroscience publications and evaluated with ground-truth annotations and scoring rubrics. b, A task example, including the task query and evaluation scoring rubrics annotated by domain experts. c, Tasks in BrainArena span diverse model organisms in neuroscience. d, BrainArena includes single- and multi-modality analysis tasks across molecule, behavior, neural activity, and structure modalities.
Fig 2: Figure 2 : BrainArena evaluation. a, Overview of BrainArena construction and evaluation process. Expert-curated tasks are generated from neuroscience publications and evaluated with ground-truth annotations and scoring rubrics. b, A task example, including the task query and evaluation scoring rubrics annotated by domain experts. c, Tasks in BrainArena span diverse model organisms in neuroscience. d, BrainArena includes single- and multi-modality analysis tasks across molecule, behavior, neural activity, and structure modalities.
Figure 4 : Structured and distributed brain-wide neural representations of behavioral types in free-moving larval zebrafish. a, Bout-level behavioral space. Each point denotes one bout ( n = 1 , 546 n=1{,}546 bouts from 18 larvae). Colors denote the three bout clusters. Faint points indicate long-duration bouts (>500 ms). Yellow markers a–i locate 9 representative skeletons in Panel b. b, Tail features of 3 bout clusters. For each cluster, the upper trace plot shows all member bouts’ tail-tip angle trajectories; the kinematic heatmap displays the mean time-varying angles of 9 tail segments from rostral (top) to caudal (bottom), using a fixed color scale spanning [ − 90 ∘ , + 90 ∘ ] [-90^{\circ},+90^{\circ}] ; the right skeleton plot exhibits the representative tail posture in different bout clusters. ros., rostral; cau. caudal. c, Region-aware spatial-functional supervoxel (SV) topology from an example larva. Each point is one of 3 , 000 3{,}000 supervoxel centroids. Colors indicate major anatomical blocks. Forebrain, orange; midbrain, green; hindbrain, blue; spinal cord, purple; ganglia, red. d, Example calcium activity heatmap of brain-wide supervoxels sorted by major anatomical block and region label. Bout-locked activity peaks are denoted at the top of the heatmap. S., spinal cord; G., ganglion. e, InfoNCE training loss of the bout type supervised CEBRA model fitted in one example larva. Dashed line, uninformative-baseline reference. f, CEBRA latent manifold for the same larva in Panel e. Gray, full-session 8D latent trajectory projected onto a 2D plotting basis derived by SVD of bout-mean latents; colored points, accepted bout-mean latents colored by bout type. g, Aggregated 3NN confusion matrix across larvae.
Fig 4: Figure 4 : Structured and distributed brain-wide neural representations of behavioral types in free-moving larval zebrafish. a, Bout-level behavioral space. Each point denotes one bout ( n = 1 , 546 n=1{,}546 bouts from 18 larvae). Colors denote the three bout clusters. Faint points indicate long-duration bouts (>500 ms). Yellow markers a–i locate 9 representative skeletons in Panel b. b, Tail features of 3 bout clusters. For each cluster, the upper trace plot shows all member bouts’ tail-tip angle trajectories; the kinematic heatmap displays the mean time-varying angles of 9 tail segments from rostral (top) to caudal (bottom), using a fixed color scale spanning [ − 90 ∘ , + 90 ∘ ] [-90^{\circ},+90^{\circ}] ; the right skeleton plot exhibits the representative tail posture in different bout clusters. ros., rostral; cau. caudal. c, Region-aware spatial-functional supervoxel (SV) topology from an example larva. Each point is one of 3 , 000 3{,}000 supervoxel centroids. Colors indicate major anatomical blocks. Forebrain, orange; midbrain, green; hindbrain, blue; spinal cord, purple; ganglia, red. d, Example calcium activity heatmap of brain-wide supervoxels sorted by major anatomical block and region label. Bout-locked activity peaks are denoted at the top of the heatmap. S., spinal cord; G., ganglion. e, InfoNCE training loss of the bout type supervised CEBRA model fitted in one example larva. Dashed line, uninformative-baseline reference. f, CEBRA latent manifold for the same larva in Panel e. Gray, full-session 8D latent trajectory projected onto a 2D plotting basis derived by SVD of bout-mean latents; colored points, accepted bout-mean latents colored by bout type. g, Aggregated 3NN confusion matrix across larvae.
Figure 6 : A shared axis of regional decoding strength and task-specific residual patterns in a mouse decision-making task. a, Pearson correlation matrix of stimulus, choice, feedback, speed, and velocity decoder maps across 201 regions. b, Spearman rank-correlation matrix for the same five decoders. c, Observed PC1 variance fraction (red line; 0.632) compared with a 5,000-permutation column-shuffle null; empirical P ≈ 0.0002 P\approx 0.0002 . d, PC1 loadings for the metric-value PCA and rank-PCA decompositions; all decoder loadings are positive. e, R 2 R^{2} from ordinary least-squares regression of each task decoder on speed and velocity. f, Main PC1 scores for the 201 included regions, rendered bilaterally and mirror-symmetrically on the Swanson mouse-brain flatmap. Higher scores occurred in reticular, motor-related brainstem, and thalamic regions. g, Rank-PCA PC1 score projected onto the same atlas, confirming robustness of the shared axis. h, Choice-decoder residual map after removing speed- and velocity-linked variance. i, Stimulus-decoder residual map after removing speed- and velocity-linked variance. j, Feedback-decoder residual map after removing speed- and velocity-linked variance. See Supplementary Table LABEL:tab:brain_region_abbr for the full names of the brain region abbreviations.
Fig 6: Figure 6 : A shared axis of regional decoding strength and task-specific residual patterns in a mouse decision-making task. a, Pearson correlation matrix of stimulus, choice, feedback, speed, and velocity decoder maps across 201 regions. b, Spearman rank-correlation matrix for the same five decoders. c, Observed PC1 variance fraction (red line; 0.632) compared with a 5,000-permutation column-shuffle null; empirical P ≈ 0.0002 P\approx 0.0002 . d, PC1 loadings for the metric-value PCA and rank-PCA decompositions; all decoder loadings are positive. e, R 2 R^{2} from ordinary least-squares regression of each task decoder on speed and velocity. f, Main PC1 scores for the 201 included regions, rendered bilaterally and mirror-symmetrically on the Swanson mouse-brain flatmap. Higher scores occurred in reticular, motor-related brainstem, and thalamic regions. g, Rank-PCA PC1 score projected onto the same atlas, confirming robustness of the shared axis. h, Choice-decoder residual map after removing speed- and velocity-linked variance. i, Stimulus-decoder residual map after removing speed- and velocity-linked variance. j, Feedback-decoder residual map after removing speed- and velocity-linked variance. See Supplementary Table LABEL:tab:brain_region_abbr for the full names of the brain region abbreviations.

Ready to evaluate your AI agents?

Learn how ReputAgent helps teams build trustworthy AI through systematic evaluation.

Learn More

Considerations

SeekBrain depends on large language models and a curated recipe library, so it can inherit model errors and requires ongoing expert curation. Evaluation used a custom benchmark and two case studies, so performance may vary on other datasets or emerging modalities such as spatial transcriptomics or large-scale connectomics. The system is researcher-guided rather than fully autonomous for high-stakes interpretation, so human validation remains necessary for definitive claims. Mutual Verification Pattern

Methodology & More

SeekBrain turns scattered neuroscience methods into an indexed, evolving library called the Neuroscience Analysis Repertoire: reusable analysis recipes extracted from papers and open-source code. A Research Planning Engine breaks a high-level question into hierarchical tasks, while an Analysis Execution Engine converts tasks into executable code, runs analyses, and validates outputs across scientific criteria. The framework adds iterative peer-review steps and supports researcher-in-the-loop edits so experts can approve or refine workflows; validated traces update the repertoire over time. On the BrainArena benchmark of 32 expert-annotated tasks, SeekBrain outperformed general coding agents (11.7% and 17.7% gains) and showed the repertoire is critical to producing valid interpretations (ablation drops >10 points). Two real-data cases illustrate practical value: guided analysis of free-moving larval zebrafish revealed compact neural manifolds tied to behavior, and autonomous reanalysis of a mouse decision-making dataset found a shared low-rank axis of regional decoder strength after separating movement effects. The design points to a scalable pattern: distill domain methods into curated recipes, use multiple specialized agents to plan and validate, and keep experts in the loop to maintain scientific rigor and grow the system’s institutional knowledge. Orchestrator-Worker Pattern Hierarchical Multi-Agent Pattern
Need expert guidance?We can help implement this
Learn More
Credibility Assessment:

Large author list but mostly low h-indices and no affiliations specified; arXiv preprint and limited high-profile signals — emerging/limited information.