Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
15 changes: 12 additions & 3 deletions .github/workflows/ci.yml
Original file line number Diff line number Diff line change
Expand Up @@ -22,9 +22,13 @@ jobs:
- name: Sync environment
run: uv sync --extra dev
- name: Format check
run: uv run ruff format --preview --check src tests
run: >-
uv run ruff format --preview --check src tests
examples/serosurvey_study.py examples/serosurvey_study.ipynb
- name: Lint
run: uv run ruff check --preview src tests
run: >-
uv run ruff check --preview src tests
examples/serosurvey_study.py examples/serosurvey_study.ipynb
- name: Type check (strict)
run: uv run mypy --strict src

Expand Down Expand Up @@ -54,13 +58,18 @@ jobs:
enable-cache: true
cache-dependency-glob: "**/pyproject.toml"
- name: Sync environment
run: uv sync --extra examples
run: uv sync --extra examples --extra notebook
- name: Run CSTR example
run: uv run python examples/cstr_study.py
- name: Run sklearn example
run: uv run python examples/sklearn_study.py
- name: Run assay cost annotation example
run: uv run python examples/assay_study.py
- name: Execute serosurvey demo notebook
run: >-
uv run --extra notebook jupyter nbconvert --execute --to notebook
--ExecutePreprocessor.timeout=60 --output serosurvey_executed.ipynb
--output-dir /tmp examples/serosurvey_study.ipynb

ci:
name: ci
Expand Down
1 change: 1 addition & 0 deletions CHANGELOG.md
Original file line number Diff line number Diff line change
Expand Up @@ -6,6 +6,7 @@ All notable changes to this project will be documented in this file.

### Added

- An executed IVAC-oriented Jupyter notebook compares synthetic serosurvey designs across cost, overall accuracy and underserved-group accuracy, with live budget/preference changes, a 15-minute presenter route and documentation figures. Install notebook tools with the `notebook` extra.
- Adaptive trial inspection, explicit failure reporting, and bounded retries with persistent failure reasons and retry lineage (#151).
- Incremental grid checkpoints preserve completed design-point/replicate evaluations across interruptions, including parallel workers and incomplete `Study` phases. `run_grid(max_retries=...)` provides opt-in bounded retries (#151).
- Grouped surrogate validation holds out whole designs or regimes alongside separate row-validation metrics. Prediction and recommendation expose observed-support diagnostics and warn on extrapolation for GP and RF (#152).
Expand Down
7 changes: 7 additions & 0 deletions README.md
Original file line number Diff line number Diff line change
Expand Up @@ -99,6 +99,12 @@ front = study.front("benchmark") # non-dominated config indices
study.compare_phases() # per-phase hypervolume and IGD+ between successive fronts
```

For a complete 15-minute demonstration with live code and saved figures, see
the [serosurvey design notebook](examples/serosurvey_study.ipynb) and
[presentation guide](https://jcm-sci.github.io/trade-study/guide/serosurvey/).
From a repository checkout, launch it with
`uv run --extra notebook jupyter lab examples/serosurvey_study.ipynb`.

### Protocols

Users implement two protocols to plug in their domain:
Expand Down Expand Up @@ -147,6 +153,7 @@ pip install trade-study[design,pareto]
| `surrogate` | scikit-learn | GP/RF score and regime surrogates |
| `dataframe` | pandas | ResultsTable export for analysis and CSV |
| `all` | All of the above | |
| `notebook` | JupyterLab, nbconvert, matplotlib, pandas, pymoo | Execute and present example notebooks |

**Core dependency**: numpy only.

Expand Down
Binary file added docs/assets/serosurvey_costs.png
Loading
Sorry, something went wrong. Reload?
Sorry, we cannot display this file.
Sorry, this file is invalid so it cannot be displayed.
Binary file added docs/assets/serosurvey_population.png
Loading
Sorry, something went wrong. Reload?
Sorry, we cannot display this file.
Sorry, this file is invalid so it cannot be displayed.
Binary file added docs/assets/serosurvey_priorities.png
Loading
Sorry, something went wrong. Reload?
Sorry, we cannot display this file.
Sorry, this file is invalid so it cannot be displayed.
Binary file added docs/assets/serosurvey_tradeoffs.png
Loading
Sorry, something went wrong. Reload?
Sorry, we cannot display this file.
Sorry, this file is invalid so it cannot be displayed.
176 changes: 176 additions & 0 deletions docs/guide/serosurvey.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,176 @@
# How much survey is enough—and whose uncertainty matters?

This notebook is a 15-minute demonstration for a mixed audience at the
Johns Hopkins International Vaccine Access Center (IVAC). It compares
serosurvey designs using financial cost, overall estimation accuracy and
accuracy for an underserved group. All populations, prevalences, prices
and preference weights are hypothetical.

Open the executed
[Jupyter notebook](https://github.com/jcm-sci/trade-study/blob/main/examples/serosurvey_study.ipynb)
to read the narrative, code, tables and saved figures. The accompanying
[Python script](https://github.com/jcm-sci/trade-study/blob/main/examples/serosurvey_study.py)
contains the simulator and plotting helpers and regenerates these figures.
Use both files from a checkout; the notebook imports the companion module.

## Run or present the notebook

From the repository root:

```bash
uv run --extra notebook jupyter lab examples/serosurvey_study.ipynb
```

Choose the environment's Python kernel, then restart it and run all cells.
The `notebook` extra supplies JupyterLab, nbconvert, pandas, matplotlib
and Pareto analysis without requiring the other optional modeling backends.
Installation needs network access; the executed example uses only local
code and synthetic data. Run from a checkout of `main`: the notebook uses
the preference API added after the 0.3.0 release.

Clean notebook executions took about six to eight seconds on the development
machine, including kernel startup and rendering. The companion script took
about two seconds for 19,800 evaluations and four figures. These measurements
exclude dependency installation; check your presentation machine before the talk.

Saved notebook outputs provide a fallback without running code. Export them
to HTML for an additional presentation copy:

```bash
uv run --extra notebook jupyter nbconvert --to html examples/serosurvey_study.ipynb
```

To verify execution in a fresh kernel without modifying the saved notebook:

```bash
uv run --extra notebook jupyter nbconvert --execute --to notebook \
--ExecutePreprocessor.timeout=60 --output serosurvey_executed.ipynb \
--output-dir /tmp examples/serosurvey_study.ipynb
```

## The decision

IVAC's [SISS project](https://publichealth.jhu.edu/ivac/our-work/strengthening-immunization-systems-through-serosurveillance-siss)
examined the design and use of serological surveillance. The
[serosurvey costing study by Carcelen, Patenaude, Moss and colleagues](https://pmc.ncbi.nlm.nih.gov/articles/PMC7561102/)
provides a concrete link between epidemiology and economic evaluation.
Its study-, cluster- and participant-level cost structure motivates this
example; we do not reproduce its study or use its historical prices.

The fictional population consists of an 80% group and a 20% underserved
group, with antibody-status prevalences of 90% and 65% respectively.
These are assumed model inputs, not measured data or protection thresholds.

![Hypothetical population assumptions](../assets/serosurvey_population.png)

Compare 18 designs:

| Factor | Levels |
|---|---|
| Participants | 300, 600, 1,200 |
| Communities | 10, 20, 40 |
| Allocation | Proportional 80:20, or oversampling 50:50 |

Allocation applies to both participants and communities. Within each group,
participants are distributed as evenly as possible, keeping exact totals.
Independent community probabilities follow a Beta distribution centered on
the group mean, with illustrative within-community correlation 0.06;
participant counts then follow a binomial model. The estimand is the fixed
group mean, not the realized mean in the sampled communities.

The overall prevalence estimator uses population weights of 80:20 for both
allocation strategies. Oversampling does not change the population composition.
Financial cost is fixed setup plus community visit costs plus participant costs.
Average per-community and per-participant costs are not added together as if
they were independent marginal costs.

## Run, aggregate and refine

The notebook visibly constructs a `Study` with two grid `Phase`s: 100 simulated
surveys per design, followed by 1,000 per design with an independent phase seed.
Both phases evaluate all 18 designs. Refinement increases simulation replication,
not the number of participants per survey. The refined estimates replace the
screening estimates; the two phases are not pooled.

Each scorer call returns financial cost and absolute prevalence errors in
percentage points. `aggregate_replicates()` averages those errors, producing
mean absolute error (MAE), and retains their Monte Carlo variation. Averaging
signed errors before taking their absolute value would measure something different.

## Inspect feasible alternatives

A `Constraint` imposes an illustrative $40,000 financial budget. The Pareto
set minimizes cost, overall MAE and underserved-group MAE simultaneously.
Subgroup error is a narrow measure of information equity, not a comprehensive
measure of equity in health outcomes.

![Cost, overall accuracy and subgroup accuracy](../assets/serosurvey_tradeoffs.png)

Gray designs exceed the budget. Outlines show the feasible Pareto set computed
using all three objectives; the two panels are projections, not independently
computed two-objective fronts. Design labels identify the preference winners.
Vertical bars are approximately two Monte Carlo standard errors of estimated
MAE. They express simulation precision under this model, not uncertainty in a
real survey's prevalence or the model assumptions. They are marginal bars,
not simultaneous post-selection confidence guarantees.

## Change priorities without rerunning simulations

The notebook passes the raw results to `preference_sweep()` with three explicit
preference vectors and `normalization="reference"`. Fixed reference ranges are
$0–70,000, 0–4 percentage points overall MAE, and 0–10 percentage points subgroup
MAE. These anchors scale preferences; they are not feasibility thresholds and
do not clip values. Scenario weights are hypothetical, not elicited stakeholder values.

![Ranks under three preference scenarios](../assets/serosurvey_priorities.png)

Rank 1 wins within a scenario. The displayed rows are feasible Pareto designs,
but ranks include all feasible designs. Choices depend on point estimates and
can change with further simulation or different assumptions. Any reported
selection fraction describes the supplied preference scenarios, not a probability
that a design is best.

For a short live interaction, change the budget to $30,000 in the decision
cell and rerun that cell and the figures below it. Restore the budget and edit
the subgroup-priority weights to compare another preference. No simulation
rerun is needed. If the budget admits no alternative, the example reports no choice.

The optional cost breakdown in the appendix supports an economics discussion:

![Cost components for the selected designs](../assets/serosurvey_costs.png)

## A 15-minute presentation

| Minutes | Content |
|---|---|
| 0–2 | The decision and the two population groups |
| 2–4 | Design factors and competing objectives |
| 4–6 | One visible `Study` definition and a live run |
| 6–10 | Budget and Pareto plots |
| 10–12 | Priorities and an optional budget change |
| 12–13 | Assumptions a real project would replace |
| 13–15 | Discussion |

The notebook includes presenter notes and slideshow cell metadata. Leave the
model details and cost breakdown as appendices for the main talk. The saved
notebook and HTML export make a Beamer build unnecessary for the current
fast-running example.

## What a real project would replace

Replace the synthetic population with context-specific prevalence, clustering,
nonresponse and sampling-frame assumptions; add validated assay characteristics
and uncertainty; use local financial and economic costs; and define objectives
and practical constraints with stakeholders. The example holds assay effects
fixed and does not equate antibody-status prevalence with complete protection.
Communities are sampled without selection bias by construction. Real selection
and nonresponse can introduce bias that more simulation cannot remove.

The source notebook links the IVAC projects and member profiles that informed
its scope. Those connections do not imply endorsement of the example.

To regenerate documentation figures:

```bash
uv run --extra notebook python examples/serosurvey_study.py
```
3 changes: 3 additions & 0 deletions docs/index.md
Original file line number Diff line number Diff line change
Expand Up @@ -7,6 +7,9 @@ sensitivity analysis, and model stacking.
For installation and quick-start examples, see the
[README](https://github.com/jcm-sci/trade-study#readme).

For a short presentation with live code, tables and saved figures, see the
[serosurvey design notebook](guide/serosurvey.md), designed for a mixed IVAC audience.

## Overview

`trade-study` provides a structured workflow for multi-objective
Expand Down
Loading
Loading