Squad composition
Why the longest work is several Expert Squads rather than one team running longer, shown through DeBERTa research engineering and an OpenCorvus self-paper Mission.
The longest work is not one team working longer. It is several teams, each owning a stage, each handing the next one something it can read.
That is a claim about failure modes, not about org charts. A single team carrying an outcome from first source to final file accumulates its own early mistakes, competes with itself for context, and ends up reviewing work it wrote. Splitting the outcome into owned stages puts a different squad — with different roles, different tools, and no stake in the previous stage’s conclusions — on each side of every handoff.
What a Mission holds
A Mission records which Expert Squad IDs are available when it starts. Capabilities installed later do not silently widen that set. Each child Task then resolves one admitted ID to one exact package revision plus its selected workflow, fixed for that Task’s lifetime.
So composition happens at the Mission level, and ownership stays at the Task level. Nothing about combining squads dilutes who is responsible for a given delivery.
The case: DeBERTa from model card to published repository
The readable chain has six complete deliveries: model and data; a CUDA-only training system with a live monitor and inference page; architecture evidence; an ACL-style short paper; independent paper review; and the Mission repository. Every count below is resolved from the catalog rather than written into this page.
6 squads44 named roles
- Acquire DeBERTa v3 Base ABSA v1.1, re-investigate current sources, and search for or synthesize traceable training data before training.
- Provision a complete CUDA-only runtime with no CPU training path; record train/test performance for every innovative design and iteration; build an auto-updating training monitor and inference website; keep improving the model against an explicit baseline.
- Use the best experiment's exact model architecture and design rationale to produce reproducible, publication-ready figures.
- Research related literature and write a complete ACL-style short paper of at least four pages.
- Deeply review and proofread the manuscript, eliminate factual errors, and make its organization concise and informative.
- Create a well-organized Git repository, execute the work through Mission decomposition, and deliver a reviewed GitHub push.
View the audited CUDA experiment artifact on GitHub ↗
| Stage | Squad | Roles | Hands on |
|---|---|---|---|
| Model & data | Deep Research | 6 | A verified DeBERTa v3 Base ABSA v1.1 source, current ABSA evidence, and a sourced plan to find, clean, or synthesize training data. |
| CUDA training system | Advanced | 14 | A CUDA-only runtime, iterative baseline and candidate training, an experiment ledger, and a live training-monitor and inference website. |
| Architecture evidence | Data Analysis & Business Insights | 7 | Best-run comparisons, architecture and design diagrams, and reproducible figures bound to the exact winning checkpoint. |
| ACL short paper | Research Studio | 5 | A concise, informative ACL-style short paper of at least four pages, grounded in related work and the best experiment. |
| Independent paper review | Academic Paper Review | 8 | Resolved findings across facts, citations, novelty, method, structure, figures, hallucination risk, concision, and informativeness. |
| Mission repository | Base | 4 | A reproducible, organized Git repository with the Mission stage map, reviewed documentation, and a verified GitHub push. |
Read the shape rather than the individual model choice. The landing page can unfold the same Mission from six deliveries into eighteen squad-owned Tasks: source and dataset lineage, CUDA runtime, training campaigns, metric reconciliation, a real monitoring and inference page, figures, related work, manuscript audit, clean-room reproduction, repository hardening, and publication. The scope gets larger without making one team author, test, review, and approve its own work.
The case: OpenCorvus researches and introduces itself
The second long Mission does not ask the system to write product copy from memory. It requires an evidence chain from the source tree, current architecture specifications, release artifacts, version history, empirical evaluation, and inspected related work, followed by a systems paper with at least 30 pages of substantive main text. References and appendices do not count, and repeated background, generic prose, oversized figures, or loose layout cannot be used as padding.
9 squads55 named roles
- Reconstruct OpenCorvus from inspected source code, current architecture specifications, documentation, releases, and version history; do not write from memory or product claims.
- Research inspected related work, define falsifiable questions and bounded contributions, then design reproducible evaluation with baselines, metrics, uncertainty, failures, and ablations.
- Use Research Studio's analysis-report-quality Skill to create one validated evidence model before drafting a systems paper with at least 30 substantive main-text pages, excluding references and appendices.
- Do not pad the paper with repeated background, generic Agent prose, oversized figures, loose spacing, appendix migration, or unsupported claims.
- Redraw every necessary figure from accepted data with publication-grade hierarchy, labels, units, sample sizes, uncertainty, provenance, captions, and accessible encodings; raw notebook or chart-library output is forbidden.
- Independently review literature, novelty, logic, methods, statistics, facts, citations, hallucination risk, and presentation; resolve every critical and major finding and deliver the manuscript, source, bibliography, evidence ledger, figure sources, and reproduction guide together.
Mission to give OpenCorvusComplete this long Mission with one complete Expert Squad owning every numbered stage. Produce a venue-neutral academic systems paper that introduces OpenCorvus and contains at least 30 pages of substantive main text, excluding references and appendices. Page count must not be inflated with repeated background, generic agent prose, oversized figures, loose spacing, appendix migration, or unsupported claims. Reconstruct the system from the current source tree, architecture specifications, documentation, release artifacts, and version history; map related multi-agent harness, long-horizon execution, evidence handoff, evaluation, and self-improvement work from inspected papers and official sources; define falsifiable research questions and bounded contributions; trace architecture, control flow, persistence, Task/Mission semantics, Expert Squad composition, evolution, permissions, recovery, and observability to exact evidence. Design and run reproducible empirical evaluations with declared baselines, datasets or cases, metrics, uncertainty, failures, ablations, and limitations; preserve runnable scripts and canonical result tables. Research Studio must use its analysis-report-quality Skill to build one validated evidence model before drafting. Every figure and table must be necessary, claim-linked, source-traceable, independently reproducible, and redrawn with conclusion-led titles, legible labels, units, sample sizes, uncertainty, captions, accessible encodings, and a deliberate publication theme. Raw notebook output, chart-library defaults, clipped labels, decorative diagrams, screenshots standing in for evidence, and visual garbage are forbidden. Draft the complete paper with abstract, introduction, research questions, related work, system design, implementation, methodology, results, ablations, discussion, threats to validity, limitations, responsible-use/security considerations, reproducibility statement, conclusion, and verified references. Then run independent literature, novelty, logic, methods/statistics/fact, citation/hallucination, and presentation reviews; resolve every critical and major finding, refine the figures again against the final prose, and deliver the manuscript, source, bibliography, evidence ledger, figure sources, and reproduction instructions together.
| Stage | Squad | Roles | Hands on |
|---|---|---|---|
| Freeze the research charter | Scientific Research Design | 4 | Falsifiable questions, bounded contributions, competing explanations, evidence needs, ethics, and a no-padding acceptance contract. |
| Build the source record | Deep Research | 6 | Inspected code, specifications, releases, history, papers, and official sources with claim-level locators and search limits. |
| Reconstruct the system | Advanced | 14 | Executable architecture evidence for control flow, persistence, Mission/Task semantics, squads, evolution, permissions, recovery, and observability. |
| Run empirical evaluation | Data Analysis & Business Insights | 7 | Reproducible baselines, cases, metrics, uncertainty, ablations, failure analysis, canonical tables, and checked quantitative claims. |
| Reproduce independently | Review & Debug | 4 | Clean-room reproduction of the claimed workflows and results, with root-cause fixes and unresolved limits recorded. |
| Position the contribution | Patent Landscape and Prior Art | 4 | A query-reproducible landscape that separates inspected prior work, overlap, defensible novelty, and unknowns. |
| Refine every figure | Office Delivery | 3 | Publication-grade architecture, experiment, ablation, and comparison figures regenerated from accepted data — never raw plotting output. |
| Write the thirty-page paper | Research Studio | 5 | One validated evidence model and a concise, information-dense 30+ page main paper whose claims, prose, tables, figures, limits, and references agree. |
| Audit and finish the manuscript | Academic Paper Review | 8 | Resolved literature, novelty, logic, method, statistics, fact, citation, hallucination, and presentation findings, followed by a final figure-to-prose reconciliation. |
The writing stage must use Research Studio’s analysis-report-quality Skill to build one validated
evidence model. Every figure is redrawn from accepted data with a conclusion-led title, units,
sample size, uncertainty, provenance, and accessible encodings; raw notebook output, chart-library
defaults, clipped labels, and decorative garbage are rejected. Academic Paper Review then audits
literature, novelty, logic, methods, statistics, facts, citation hallucinations, and presentation,
and every critical or major finding must be resolved.
Other combinations that already ship
The same shape, in three shorter domains:
| Other combinations | Squad | Roles |
|---|---|---|
| Deal due diligence | Mergers and Acquisitions Due DiligenceForensic Accounting InvestigationsCommercial LegalTax ComplianceInternal Audit Control Assurance | 29 |
| Incident to written knowledge | Service Reliability Incident OperationsDigital Forensics Incident InvestigationReview & DebugKnowledge Base Operations | 18 |
| Launching something | Product ManagementMarketing & Growth StrategySEO & Generative Engine OptimizationProduct Video ProductionLocalization & Adaptation | 26 |
How to choose the split
Split where a delivery can be independently owned, accepted, or depended on. Splitting for its own sake produces coordination overhead with no owner.
In practice that means:
- A new stage needs a different kind of evidence. Sources, transactions, rendered pages, and statistical results are not checked the same way, and the squad that checks one is not the squad that checks another.
- A new stage needs an independent reader. If the point of the stage is to disagree with the previous one, it has to be a different squad — the same roster reviewing its own output is a review in name only.
- A new stage produces something downstream work depends on. A typed Artifact with provenance is what makes the dependency real rather than narrative.
Everything else belongs inside the squad you already have.
Related
- Where long-horizon work breaks — the three failures composition addresses
- How Missions run long-horizon work — dependency graph, freezing, and reopening
- Expert Squad market — every squad named above, with its roles and workflows
- How squads evolve — revising a squad in the chain