Skip to content

Latest commit

 

History

56 Commits

Folders and files

NameName
Last commit message
Last commit date
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 

Repository files navigation

CAi Copilot

CAi Copilot

An open-source agent for drug design—from research intent to executable workflows and traceable molecular candidate evidence

💻 Explore the open-source code · 🤗 Discover the Daily Paper · 📄 Read the full preprint

GitHub project homepage Hugging Face Daily Paper arXiv 2608.06961

Try CAi Copilot on Intern·Inkstone Python 3.11+ Apache-2.0 License 30+ models and tools

English | 简体中文


✨ CAi in One Minute

CAi Copilot turns open-ended drug-design requests into executable and traceable agentic workflows. Researchers can describe a target, lead structure, design constraints, and evaluation objectives in natural language. CAi then orchestrates molecular generation, structure validation, property prediction, target-specific evaluation, multi-objective screening, and iterative optimization while preserving candidate molecules, supporting evidence, run artifacts, and execution records.

CAi does more than return a set of molecules. It organizes the complete design process and delivers candidate evidence that can be screened, iterated, and reviewed.

Three Core Technical Foundations

Technical capability Value
🧭 Drug-design-specific Agent Harness Interprets research goals and constraints, plans tasks, orchestrates tools, handles runtime feedback, and tracks results
🧰 30+ molecular model and scientific tool adapters Covers molecular generation, ADMET, synthesizability, structure prediction, molecular docking, and simulation
🔁 Generate–evaluate–screen–redesign loop Optimizes candidates from intermediate results and retains the most valuable options within defined constraints and compute budgets

🚀 Online Demo

Tip

CAi Copilot is now available through the Intern·Inkstone Scientific Application Hub. You can try it online without first deploying the local toolchain.

After opening the application hub, locate CAi Copilot Molecular Design Assistant, follow the platform instructions to sign in, and enter the workspace. Application visibility and authentication may change with the deployment policy; please follow the instructions shown on the platform.


🧬 Core Capabilities

Research-intent-driven workflows

CAi accepts natural-language tasks together with protein, molecular, and other input files. It decomposes research objectives into dependency-aware execution steps and continuously consolidates results throughout the run.

flowchart LR
    A[Research goals and constraints] --> B[Task interpretation and planning]
    B --> C[Molecule generation]
    C --> D[Structure and property evaluation]
    D --> E[Multi-objective screening]
    E --> F{Goals satisfied?}
    F -- No --> C
    F -- Yes --> G[Candidate molecules and evidence]
    G --> H[Expert review and downstream experiments]
Loading

Molecular design outputs ready for downstream use

CAi produces more than natural-language answers. It also provides connected, downloadable, and structured outputs:

  • Candidate molecular structures and SMILES
  • Property, toxicity, synthesizability, and target-related computational results
  • Multi-objective screening criteria, ranking rationale, and retention reasons
  • Molecular files, result tables, and computational artifacts
  • Tool execution status, error messages, and execution records

Researchers can select candidates for another round of analogue expansion or multi-objective optimization, then advance them toward chemical synthesis and experimental validation after synthesizability assessment and expert review.

Scientific tools in isolated sandboxes

Each scientific computation runs in an isolated working directory. Tool inputs, standard output, errors, and result files remain associated with the corresponding task, making it easier to diagnose issues, reproduce runs, and integrate additional models.


🧰 Model and Tool Ecosystem

CAi provides more than 30 model and tool adapters. Each capability can be called independently or composed by the agent into an end-to-end workflow.

Task area Representative capabilities
Molecular generation and optimization RXNFlow, REINVENT 4, LibINVENT, DrugEx, DeepChem, Scaffold, SC2Mol
Structure and target evaluation Boltz, AutoDock Vina, protein and ligand preparation, interaction analysis
Molecular properties and ADMET ADMET-AI, solubility, lipophilicity, absorption, metabolism, clearance, and multiple toxicity endpoints
Synthesis-related evaluation SCScore, SAScore, RAScore, ASKCOS, SynLlama
Simulation and free-energy calculation GROMACS, FEP workflows, and result analysis
Candidate prioritization Multi-objective constraints, candidate ranking, best-so-far tracking, and iterative comparison

Note

This repository provides unified invocation interfaces and tool adapters. Some tools require separate Conda environments, model-weight downloads, external service configuration, or GPUs. Use installation reports and health-check results to determine which capabilities are available in a given deployment.


💬 Example Tasks

Case 1: Target-based de novo molecular design
Use HIV-1 protease as the target protein and 1HVR.pdb as the target structure.
Set the binding-site center to [15.2, 23.5, 6.8].
Use suitable generation tools to propose candidate small molecules, perform the
necessary structure checks, and organize the candidates and associated files
according to the Vina results.

Capabilities demonstrated: goal interpretation, protein and pocket inputs, candidate generation, molecular docking, result ranking, and file export.

Case 2: Analogue expansion around a lead scaffold
Generate 10 structurally reasonable analogues from the supplied core scaffold
while retaining the specified core. Evaluate candidate synthesizability and
basic properties, then prioritize the candidates according to the results.

Capabilities demonstrated: scaffold-constrained generation, R-group expansion, structure validation, synthesis-related evaluation, and candidate screening.

Case 3: Multi-objective candidate optimization
Analyze the key properties of the available candidate molecules and screen them
using the specified hard constraints and priority objectives. If the candidates
do not meet the objectives, continue generation or optimization while retaining
the changes and best result from each iteration.

Capabilities demonstrated: multi-objective constraints, result-driven redesign, best-so-far retention, and process traceability.


⚡ Local Deployment Quickstart

Requirements

  • Python 3.11+
  • Conda or a compatible environment manager
  • Access to an LLM API or a local inference service compatible with the OpenAI API
  • Separate model weights, external services, or GPUs for some scientific tools

1. Clone and install CAi

git clone https://github.com/datamllab/CAi_copilot.git
cd CAi_copilot

conda create -n CAi python=3.11
conda activate CAi
pip install -e .

2. Configure the model

Add the model configuration to CAi/.env:

LLM_MODEL=your-model-name
LLM_API_KEY=your-api-key

# For a custom service compatible with the OpenAI API
# LLM_BASE_URL=http://your-endpoint/v1/

TOOL_SERVER_HOST=0.0.0.0
TOOL_SERVER_PORT=8001

3. Install the required scientific tools

Tools use mutually isolated runtime environments. We recommend installing only the tools needed by your use case instead of installing every environment at once:

cd CAi/toolkit/server
bash install_all.sh vina scscore toxicity
cd ../../..

Some model source code or weights must be downloaded separately. To integrate a new tool, see the backend tool development guide. For production deployment, see the deployment guide.

4. Start the tool service and agent

Open two terminals and ensure that both use the same CAi configuration:

# Terminal 1: tool service
python -m CAi.toolkit.server.app
# Terminal 2: Web UI
python CAi/main.py

# Or use the CLI
python CAi/main.py --cli

After the tool service starts, verify the tools that were actually loaded through the health-check endpoint:

curl http://localhost:8001/health

🧱 System Architecture

Research Interface
    │  Natural-language goals, structure files, constraints, and preferences
    ▼
Agent Harness
    │  Task planning, Skills, execution loop, error handling, and result tracking
    ▼
Scientific Tool Server
    │  Unified interfaces, isolated environments, task sandboxes, and compute scheduling
    ▼
Models & Tools
    │  Generation, properties, structures, docking, simulation, and synthesis-related evaluation
    ▼
Candidate Evidence
       Candidate molecules, metrics, files, ranking rationale, and execution records

For implementation details, see the execution mechanism, Web UI backend, and deployment guide.

View the project structure
CAi_copilot/
├── CAi/
│   ├── main.py                  # Web UI / CLI entry point
│   ├── CAi_agent/               # Agent Harness, Skills, and execution system
│   ├── toolkit/                 # Agent-facing scientific tool interfaces
│   │   ├── functions/           # Molecular generation, evaluation, and optimization wrappers
│   │   └── server/              # FastAPI tool service and sandboxed tasks
│   └── web_ui/                  # Web frontend and backend
├── agent_workspace/             # Session artifacts and reusable utilities
├── benchmarking/                # Evaluation and comparison tasks
├── docs/                        # Architecture and development documentation
└── tests/                       # Automated tests

🧩 Extend CAi

Add a scientific tool

  1. Configure an isolated runtime environment and execution script under CAi/toolkit/server/tools/<tool_name>/.
  2. Add an Agent wrapper with input validation under CAi/toolkit/functions/.
  3. Export the new function and reload the tool catalog.
  4. Perform a real smoke test through /health, a direct HTTP call, and the Agent workflow.

See the backend tool development guide for detailed instructions.

Add a Skill

Add a Markdown file containing the task instructions and metadata to CAi/CAi_agent/skills/. Skills preserve validated workflows and operating procedures and are loaded only when needed.


🗺️ Roadmap

  • Improve long-running workflow execution, failure recovery, and asynchronous task management
  • Expand validated ADMET, structure-prediction, simulation, and free-energy workflows
  • Improve visualization of candidate evidence, ranking rationale, and computational provenance
  • Strengthen multi-user task isolation, queue scheduling, and resource management
  • Build more reproducible, real-world molecular design cases and benchmarks

🤝 Contributing

Contributions of molecular design models, scientific tools, Skills, test cases, and documentation improvements are welcome.

When submitting a contribution, please also describe:

  • The problem addressed by the tool or workflow
  • Required environments, model weights, and compute resources
  • Input/output contracts and error-handling behavior
  • A minimal reproducible test case
  • The applicable scope and scientific limitations of the results

⚠️ Scope and Limitations

CAi Copilot is intended for research and engineering exploration. Generated molecules, property predictions, docking results, free-energy calculations, and synthesis-related scores are computational evidence. They do not replace medicinal-chemistry expertise, experimental validation, safety assessment, or clinical decision-making. Use CAi only in accordance with applicable data licenses, model licenses, and institutional policies.


📄 Citation

If CAi Copilot contributes to your research or development, please consider citing:

@misc{cai_molecule_design_copilot_2026,
  author    = {XXX},
  title     = {CAi Molecule Design Copilot},
  year      = {2026},
  month     = {May},
  publisher = {GitHub},
  note      = {An agentic platform for molecular generation, evaluation, and candidate selection}
}
Start from research intent. Orchestrate tools. Preserve evidence. Keep designing.

About

CAi Molecule Design Copilot

Resources

Stars

31 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages