diff --git a/README.md b/README.md index 751e166..4159ef1 100644 --- a/README.md +++ b/README.md @@ -1,126 +1,76 @@ -# policy-effect-analytics-agent +# 정책 효과 분석 플랫폼 · policy-effect-analytics-agent -> **에이전틱 AI × 데이터 — 문제 정의부터 효과 분석 자동화까지** -> NIPA OpenUp 오픈소스 AI 특화형 2차 · Track 3 · 멘토 신진수 (가짜연구소 인과추론팀) +**사람들이 묻는 정책, 정말 효과가 있었을까?** +소셜 반응에서 출발해, 그 주제의 정책을 전부 모으고, 공공데이터로 효과를 추정하는 오픈소스입니다. +결론을 낼 수 없으면 "식별 불가"라고 말하는 것까지가 이 프로젝트의 일입니다. -**English summary.** An open-source toolkit and case library for estimating the effects of Korean public policies with open data. Each group writes a pre-registered `plan.yaml`, fetches public data, runs a causal estimator (DiD / event study / synthetic control …), and publishes a reproducible report. LLM agents (LangGraph + open LLMs) automate collect → metrics → estimate → report. Results are browsable in a Streamlit app. +[대시보드 보기](https://causalinferencelab.github.io/policy-effect-analytics-agent/) · [어떻게 동작하나](https://causalinferencelab.github.io/policy-effect-analytics-agent/architecture.html) · [조별 운영 가이드](docs/ops/group-guide.md) · [GitHub 처음이라면](docs/ops/github-onboarding.md) ---- - -**대시보드**: https://causalinferencelab.github.io/policy-effect-analytics-agent/ (GitHub Pages, `main` 반영 시 자동 배포) - -## 왜 하나요? +> 가짜연구소 인과추론팀 × NIPA 오픈업 오픈소스 AI 특화형 2차 트랙3 「에이전틱 AI × 데이터」 -1. 공공·사회 문제를 **데이터로 정의**하고, 그 해법(정책)이 **실제로 어떤 효과를 냈는지 추정**합니다. -2. 수집 → 지표 구조화 → 효과 추정 → 리포팅 전 과정을 **LLM 에이전트로 자동화**합니다. -3. 같은 문제를 다루는 누구나 바로 쓸 수 있도록 **GitHub + 분석 플랫폼(Streamlit)** 으로 공개합니다. - -핵심 원칙: **결과를 보기 전에 `plan.yaml`을 먼저 커밋한다.** (사전 등록 → 사후 끼워맞추기 방지) +--- -## 저장소 구조 +## 한눈에 보기 ``` -core/ 공통 엔진 (Data 담당) — 어댑터, schema/plan.py, estimators, report, agent -cases/ - _template/ 새 케이스 시작용 템플릿 (복사해서 사용) - _example_*/ 참고용 예시 케이스 - <조-주제>/ 조별 케이스 (plan.yaml, fetch.py, estimate.py, report.md, figures/) -app/streamlit_app.py 케이스 브라우저 (API 키 없이 실행) -docs/ops/ 조별 운영 가이드, GitHub 온보딩, 모니터링 -catalog/ 정책 × 데이터셋 카탈로그 (이슈 → 데이터 연결) -docs/strategy/ 문제 정의·전략 문서 -scripts/ 운영 스크립트 (weekly_activity.py 등) -tests/ 테스트 -.github/ CI, PR/이슈 템플릿, CODEOWNERS +소셜 신호 ─▶ 주제 ─▶ 정책 전부 모으기 ─▶ 세 관문 ─▶ 계획 먼저 ─▶ 효과 추정·판정 +(뉴스·SNS) (신호는 여기까지만) (법제처 조례·고시) (언제·누가·무엇을) (plan.yaml 커밋) (식별됨·조건부·식별 불가) ``` -## 빠른 시작 +| 원칙 | 왜 | +|---|---| +| 소셜 신호는 **주제까지만** 정한다 | 화제가 된 정책 하나만 고르면 결과를 보고 사례를 고르는 셈이 된다 | +| 계획을 **먼저 커밋**해야 추정이 실행된다 | 결과를 본 뒤 설계를 바꾸는 것을 막는다 | +| 방법은 **규칙이** 고르고, 수치는 **라이브러리가** 계산한다 | LLM은 주제 매칭과 서술만 돕는다 | +| 처치 지역이 적으면 **무작위화 추론** | 기존 표준오차는 처치 4/25곳에서 크게 과신한다 | +| 한계는 **배지로 공개** | 시뮬레이션·키 대기·확인 필요 상태를 숨기지 않는다 | + +## 5분 만에 돌려 보기 ```bash git clone https://github.com/CausalInferenceLab/policy-effect-analytics-agent.git cd policy-effect-analytics-agent +python -m venv .venv && source .venv/bin/activate +pip install -e ".[dev]" -# uv 권장 (pip도 가능: python -m venv .venv && pip install -e ".[dev]") -uv venv -p 3.11 && source .venv/bin/activate -make install # = uv pip install -e ".[dev]" -cp .env.example .env # API 키 입력 (공공데이터포털, KOSIS, LLM) - -make check # ruff + pytest -make app # Streamlit 케이스 브라우저 (http://localhost:8501) +make flow CASE=cases/t3-land-permit-2025 # 토허구역 예시를 6단계로 실행 +python site/build.py && python -m http.server -d _site # 대시보드를 http://localhost:8000 에서 ``` -에이전트/인과 추가 기능: `uv pip install -e ".[agent,causal]"` - -## 케이스 추가하기 (4단계) - -```bash -git switch -c group3/plan -cp -r cases/_template cases/group3-youth-rent # 폴더명: <조>-<주제>, 소문자-하이픈 -``` +API 키 없이 돌아갑니다. 키가 필요한 데이터는 `.env.example`을 복사해 채우세요(`.env`는 커밋되지 않습니다). -1. **질문 정의** — `plan.yaml` 작성 (질문·처치·대조·시점·지표·추정법·가정·중단조건·데이터 라이선스) → **먼저 PR** -2. **수집** — `fetch.py`: 공공데이터 → `data/raw/`(커밋 금지) → 정제 결과만 `data/processed/` -3. **추정** — `estimate.py`: `core.estimators`로 효과 추정 + 반증(placebo 등) → `figures/*.png` -4. **리포트** — `report.md`: 결과·한계·정책 시사점. `make app`에서 바로 보입니다. +## 무엇이 들어 있나 -자세한 절차: [`cases/_template/README.md`](cases/_template/README.md), 협업 규칙: [`CONTRIBUTING.md`](CONTRIBUTING.md), GitHub가 처음이라면: [`docs/ops/github-onboarding.md`](docs/ops/github-onboarding.md), 조별 운영: [`docs/ops/group-guide.md`](docs/ops/group-guide.md) - -## 7주 로드맵 - -| 주차 | 목표 | 산출물 (커밋 기준) | +| 폴더 | 하는 일 | 조원 역할 | |---|---|---| -| 1 | 온보딩·조 편성·주제 후보 | 이슈 `케이스 제안` 등록, 첫 PR(자기소개/브랜치) | -| 2 | 문제 정의·데이터 탐색 | `cases/<조>/plan.yaml` 초안 PR (결과 보기 전) | -| 3 | 수집 자동화 | `fetch.py`, 데이터 출처·라이선스 명시 | -| 4 | 지표 구조화·1차 추정 | `estimate.py`, 기본 그림 | -| 5 | 강건성·반증 + 에이전트화 | placebo/민감도, LangGraph 노드 연결 | -| 6 | 리포트·플랫폼 | `report.md`, Streamlit 반영 | -| 7 | 발표·회고·공개 정리 | 최종 PR 머지, 릴리스 태그 | - -## 이슈 → 주제 → 정책 전체 → 효과 - -소셜 반응(뉴스 제목·SNS 글)은 **어느 주제를 볼지까지만** 정합니다. 화제가 된 정책 하나만 골라 분석하면 -결과를 보고 사례를 고르는 셈이 되기 때문입니다(출발 키트 05). 주제가 정해지면 그 주제의 정책을 -전부 모으고(법제처 조례·고시), 세 관문(언제·누가·무엇을)을 통과한 것만 분석합니다. -주제 목록: `catalog/topics.yaml` · 정적 대시보드: `python site/build.py` → `_site/` - -1. **정책 식별**: `catalog/policies.yaml`(정책 × 데이터셋 카탈로그)에서 키워드로 찾습니다. LLM은 후보 중에서 고르는 보조 역할만 합니다. -2. **데이터셋 추천**: 카탈로그에 검증해 둔 데이터셋과 공공데이터포털 실시간 검색 결과를 보여줍니다. -3. **분석 설계**: 카탈로그의 설계(이중차분·합성통제·단절 시계열)로 정합니다. LLM이 고르지 않습니다. -4. **효과 분석**: 사전 등록된 `plan.yaml`로 아래 Flow를 실행합니다. - -```bash -make app # 사이드바 '이슈 → 데이터 → 효과' -python -c "from core.discovery import discover; r=discover('토허제 확대하고 집값 잡혔나'); print(r.top.name, r.next_step)" -``` +| [`catalog/`](catalog/) | 주제 6개 · 정책 8개 · 추천 데이터셋 | 문제 정의 | +| [`core/discovery/`](core/discovery/) | 소셜 신호 → 주제 → 정책 목록 | 문제 정의 | +| [`core/adapters/`](core/adapters/) | 국토부 실거래가 · KOSIS · 법제처 · 파일 수집 | 데이터 수집 | +| [`core/estimators/`](core/estimators/) | DiD · 이벤트 스터디 · ITS · 무작위화 추론 · 판정 | 추정 | +| [`core/agent/`](core/agent/) | 6단계 Flow, 사전 등록 게이트, 과잉해석 가드 | 리포트·에이전트 | +| [`cases/`](cases/) | 조별 분석 케이스 (`_template`에서 시작) | 조 전체 | +| [`site/`](site/) | 공개 대시보드 (GitHub Pages) | 리포트·에이전트 | +| [`app/`](app/) | Streamlit 개발용 화면 | — | +| [`docs/`](docs/) | 전략(국내 사례·주제 가이드·계획 작성법), 운영 가이드 | — | -샘플: [`cases/t3-land-permit-2025`](cases/t3-land-permit-2025/) (현재 시뮬레이션 데이터, API 키 발급 후 실데이터로 전환) +## 조별로 참여하기 -## Flow — 6단계 에이전트 흐름 +1. **주제 고르기**: [대시보드](https://causalinferencelab.github.io/policy-effect-analytics-agent/#topics)에서 주제를 고르거나 `catalog/topics.yaml`에 제안합니다. +2. **계획 PR**: `cp -r cases/_template cases/group1-<주제>` → `plan.yaml` 작성 → PR로 사전 등록합니다. 작성법은 [plan-guide](docs/strategy/plan-guide.md)를 보세요. +3. **실행·공개**: `make flow CASE=cases/group1-<주제>`를 돌리고 결과를 PR로 올리면, `main`에 반영될 때 대시보드에 자동으로 올라갑니다. -`core.agent`가 케이스 하나를 아래 6단계로 실행하고 `cases/<케이스>/flow_log.json`에 단계별 결과를 남깁니다. 앱의 **Flow** 페이지(`make app` → 사이드바 Flow)에서 실행하거나 저장된 로그를 볼 수 있습니다. +브랜치는 `group/<설명>`, 수정은 자기 조 폴더만, 병합은 리뷰 1명 + CI 통과 후입니다. 자세한 규칙은 [CONTRIBUTING](CONTRIBUTING.md)에 있습니다. -| 단계 | 하는 일 | 주차 | -|---|---|---| -| ① 문제 정의 | `plan.yaml` 검증 + 사전 등록(커밋) 확인 — 추정 전에 확정 | W3 | -| ② 데이터 수집 | 공공데이터 수집·출처/라이선스 기록 | W2 | -| ③ 지표 구조화 | 패널 구성·품질 점검 | W3 | -| ④ 효과 추정 | DiD/이벤트 스터디/ITS + 반증 → 식별됨·조건부·식별 불가 | W4–5 | -| ⑤ 과잉해석 가드 | 결론 보류 규칙 + 인과 단정 표현 검사 | W6 | -| ⑥ 리포트 | `report.md`·그림·재현 기록 | W7 | +## 지금 상태 -```bash -make flow CASE=cases/_example_night_clinic # = python -m core.agent <케이스> --allow-uncommitted -``` - -`--allow-uncommitted`는 데모용입니다. 실제 분석은 `plan.yaml`을 먼저 커밋한 뒤 플래그 없이 실행하세요. LLM 서술은 `.env`에 키를 넣고 `--llm`으로 켭니다(`uv pip install -e ".[agent]"`). +- **구현됨**: 주제 매칭 · 6단계 Flow · 사전 등록 게이트 · DiD/이벤트 스터디/ITS · 무작위화 추론 · 과장 표현 가드 · 대시보드 자동 배포(매주 월요일 갱신) +- **키 대기**: 국토부 실거래가(토허구역 예시는 지금 **시뮬레이션 데이터**) · 법제처 조례 자동 수집 +- **다음 단계**: 시차 도입 추정(Callaway–Sant'Anna) · 합성통제 · 검색량으로 선반영 점검 ## 라이선스 -- **코드**: MIT ([LICENSE](LICENSE)) — © 가짜연구소 Causal Inference Team -- **데이터**: 각 출처의 이용 조건을 따릅니다 (공공누리 제1~4유형, KOSIS 이용약관 등). 각 케이스의 `plan.yaml > data_sources[].license`에 반드시 명시하고, 재배포가 제한된 원자료는 커밋하지 않습니다. -- 개인정보가 포함된 원자료는 어떤 경우에도 커밋하지 않습니다. +코드는 [MIT](LICENSE)입니다. 데이터는 각 출처의 이용 조건(공공누리 유형 등)을 따르며, 케이스마다 `plan.yaml`의 `data_sources`에 적습니다. -## 기여 +--- -[CONTRIBUTING.md](CONTRIBUTING.md) · [CODE_OF_CONDUCT.md](CODE_OF_CONDUCT.md) +**English.** An open-source platform that starts from public conversation, picks a *topic* (never a single trending policy, to avoid selecting cases on outcomes), collects every policy in that topic from Korean public sources, checks three gates (when / who / what), requires a committed pre-analysis plan, and estimates effects with rule-selected designs (DiD, event study, ITS; randomization inference when few units are treated). Results are published as a static dashboard on GitHub Pages. diff --git a/site/architecture.py b/site/architecture.py new file mode 100644 index 0000000..7fde971 --- /dev/null +++ b/site/architecture.py @@ -0,0 +1,182 @@ +"""아키텍처 페이지 본문 — 지침에서 출발해 어떤 판단을 거쳐 지금 구조가 됐는지 순서대로 설명한다.""" + +from __future__ import annotations + +REPO = "https://github.com/CausalInferenceLab/policy-effect-analytics-agent" + +GOALS = [ + ("문제를 데이터로 정의하고 효과를 추정", "공공·사회 문제의 해법이 실제로 어떤 효과를 냈는지"), + ("전 과정을 LLM 에이전트로 자동화", "수집 → 지표 구조화 → 효과 추정 → 리포팅"), + ("누구나 바로 쓰는 오픈소스", "GitHub와 분석 플랫폼으로 공유"), +] + +# (질문, 근거, 결정, 어디에 반영됐나) +DECISIONS = [ + ( + "소셜 신호로 무엇을 정하나?", + "출발 키트 05: 화제성으로 사례를 고르면 결과를 보고 고르는 셈 → 효과 과대 추정", + "신호는 주제까지만 정한다. 분석은 그 주제의 정책 전체로", + "catalog/topics.yaml · core/discovery", + ), + ( + "정책 목록은 어떻게 모으나?", + "출발 키트 05: 법제처 조례 API는 무료·즉시 승인, '어느 지역이 언제부터'가 한 줄에", + "조례형 주제는 법제처 API로 자동 수집, 고시형 주제는 에이전트가 고시문에서 이력 수집", + "core/adapters/law.py · 주간 Actions", + ), + ( + "분석해도 되는 주제인지 어떻게 거르나?", + "출발 키트의 세 관문: 언제 시작했나 · 누가 받았나 · 무엇으로 재나", + "주제마다 관문 상태를 기록하고, 통과 못 하면 추정하지 않는다", + "topics.yaml gates · 대시보드 배지", + ), + ( + "LLM에게 어디까지 맡기나?", + "목표 2(자동화) vs 인과 판단의 과장 위험. 선행 시스템 CAIS도 규칙 기반 방법 선택을 택함", + "LLM은 주제 매칭 보조·서술만. 추정 방법은 데이터 모양으로 규칙이 고르고, 수치는 라이브러리가 계산", + "core/estimators · core/agent/guard.py", + ), + ( + "결과를 보고 계획을 바꾸면?", + "공공 평가 기관의 사전 등록 관행(GSA OES, World Bank DIME)", + "plan.yaml을 git에 커밋해야 추정 단계가 실행된다", + "core/agent/nodes.py ① 게이트", + ), + ( + "처치 지역이 몇 곳뿐이면?", + "토허구역 사례: 25개 구 중 4개. 기존 표준오차는 사전추세가 없어도 50~80% 기각", + "처치 단위 10곳 미만이면 무작위화 추론으로 자동 전환", + "core/estimators/ri.py", + ), + ( + "누구나 보려면 어디에 공개하나?", + "출발 키트 04: Streamlit은 서버가 필요, 서버 없이 GitHub Pages로도 가능", + "정적 대시보드를 Actions가 만들어 Pages로 배포. Streamlit은 개발용", + "site/ · .github/workflows/pages.yml", + ), +] + +LAYERS = [ + ( + "인터페이스", + "누가 보나", + [ + ("대시보드", "GitHub Pages · 누구나"), + ("Streamlit 앱", "개발·디버깅용"), + ("CLI", "make flow CASE=…"), + ], + ), + ( + "에이전트", + "순서대로 실행", + [ + ("① 문제 정의", "plan.yaml 검증 + 사전 등록 게이트"), + ("② 수집 · ③ 지표", "어댑터 실행, 품질 점검"), + ("④ 추정", "규칙이 고른 방법으로"), + ("⑤ 가드 · ⑥ 리포트", "과장 차단, 재현 기록"), + ], + ), + ( + "분석 코어", + "LLM 없음, 결정론", + [ + ("분석계획 스키마", "core/schema"), + ("추정기", "DiD · 이벤트 스터디 · ITS"), + ("추론", "무작위화 추론(소수 처치)"), + ("판정", "식별됨 · 조건부 · 식별 불가"), + ], + ), + ( + "데이터", + "무엇을 모으나", + [ + ("카탈로그", "주제 6 · 정책 8 · 데이터셋"), + ("어댑터", "국토부 · KOSIS · 법제처 · 파일"), + ("스냅샷", "출처·라이선스·해시 기록"), + ], + ), +] + +CODE_MAP = [ + ("catalog/", "주제·정책·데이터셋 카탈로그", "문제 정의", "W2–W3"), + ("core/discovery/", "소셜 신호 → 주제 → 정책 목록", "문제 정의", "W3"), + ("core/adapters/", "공공데이터·법제처 수집", "데이터 수집", "W2"), + ("core/schema/", "분석계획(plan.yaml) 규칙", "문제 정의", "W3"), + ("core/estimators/", "추정·검증·판정", "추정", "W4–W5"), + ("core/agent/", "6단계 Flow, 과잉해석 가드", "리포트·에이전트", "W5–W6"), + ("cases/<조>/", "조별 분석 케이스", "조 전체", "W3–W7"), + ("site/", "공개 대시보드", "리포트·에이전트", "W7"), +] + +STATUS = [ + ( + "go", + "구현됨", + "주제 매칭 · 6단계 Flow · 사전 등록 게이트 · 이벤트 스터디/DiD/ITS · 무작위화 추론 · 과장 표현 가드 · 대시보드 자동 배포", + ), + ( + "warn", + "키 대기", + "국토부 실거래가(실데이터 전환) · 법제처 조례 자동 수집(시크릿 등록 후 다음 배포부터)", + ), + ( + "info", + "다음 단계", + "시차 도입 추정(Callaway–Sant'Anna, W4) · 합성통제(W5) · 검색량으로 선반영 점검(데이터랩 키)", + ), +] + + +def body(e) -> str: + goals = "".join( + f'
{i + 1}

{e(a)}

{e(b)}

' + for i, (a, b) in enumerate(GOALS) + ) + decisions = "".join( + f'
' + f"결정 {i + 1:02d}

{q}

" + f'

근거 — {why}

→ {what}

' + f'

{where}

' + for i, (q, why, what, where) in enumerate(DECISIONS) + ) + layers = "".join( + f'
' + f'

{name}

{hint}
' + f'
' + + "".join( + f'
{a}' + f'
{b}
' + for a, b in items + ) + + "
" + for name, hint, items in LAYERS + ) + code = "".join( + f"/', '')}'>{e(p)}" + f"{e(r)}{e(w)}{e(wk)}" + for p, r, w, wk in CODE_MAP + ) + status = "".join( + f'
{lab}

{e(t)}

' + for c, lab, t in STATUS + ) + return f""" +
ARCHITECTURE +

어떻게 동작하나

+

과정 목표 세 가지에서 출발해, 어떤 질문에 어떤 근거로 답했는지 순서대로 적었습니다. +지금 구조는 그 답들을 쌓은 결과입니다.

+ +

출발점: 과정 목표

{goals}
+ +

결정의 흐름

목표를 코드로 옮기면서 부딪힌 질문 일곱 개입니다.

+
{decisions}
+ +

네 개의 층

위에서 아래로 호출합니다. LLM은 에이전트 층에서 서술과 주제 매칭만 돕고, 분석 코어에는 들어가지 않습니다.

+
{layers}
+

옆에서 GitHub Actions가 돕습니다: 테스트(CI) · 매주 데이터 갱신 · 대시보드 배포.

+ +

코드 지도

폴더마다 맡는 조원 역할과 주차입니다.

+
{code}
폴더하는 일조원 역할주차
+ +

지금 상태

{status}
+""" diff --git a/site/build.py b/site/build.py index 3939c1f..ce481c7 100644 --- a/site/build.py +++ b/site/build.py @@ -46,46 +46,65 @@ e = html.escape CSS = """ -:root{--paper:#f4f5f7;--surface:#fff;--ink:#16181d;--ink2:#454b57;--muted:#6f7683;--hair:#d9dce2; ---accent:#1d5b8f;--go:#1f6b4f;--go-bg:#e2f1ea;--warn:#8a5a00;--warn-bg:#fbefd7;--stop:#9b2c33;--stop-bg:#f8e3e4; ---sans:"Pretendard","Apple SD Gothic Neo","Malgun Gothic",system-ui,sans-serif} -@media (prefers-color-scheme:dark){:root{--paper:#121418;--surface:#1b1e24;--ink:#eceef2;--ink2:#c3c8d1;--muted:#8d94a1; ---hair:#2e333c;--accent:#7cb4e6;--go:#79d1a8;--go-bg:#16302a;--warn:#e7b35a;--warn-bg:#33291a;--stop:#ec8e95;--stop-bg:#361e21}} -*{box-sizing:border-box}body{margin:0;background:var(--paper);color:var(--ink);font:16px/1.7 var(--sans);word-break:keep-all} -a{color:var(--accent)}.wrap{max-width:1040px;margin:0 auto;padding:40px 16px 80px} -header.top{display:flex;justify-content:space-between;align-items:baseline;gap:12px;flex-wrap:wrap;border-bottom:2px solid var(--ink);padding-bottom:14px} -header.top a{font-size:14px}h1{font-size:clamp(28px,4.4vw,40px);line-height:1.2;margin:22px 0 8px} -h2{font-size:21px;margin:44px 0 12px}h3{font-size:17px;margin:0 0 6px}.lede{color:var(--ink2);max-width:66ch;margin:0} -.muted{color:var(--muted);font-size:14px}.card{background:var(--surface);border:1px solid var(--hair);border-radius:10px;padding:18px} -.grid{display:grid;gap:14px;grid-template-columns:repeat(auto-fit,minmax(290px,1fr))} -.steps{display:grid;gap:10px;grid-template-columns:repeat(auto-fit,minmax(170px,1fr));margin-top:16px} -.step{background:var(--surface);border:1px solid var(--hair);border-radius:10px;padding:12px 14px;font-size:14px} -.step b{display:block;font-size:15px}.step .n{color:var(--accent);font-weight:700;font-size:13px} +:root{--paper:#f6f7f9;--surface:#fff;--sunk:#eef0f3;--ink:#14171c;--ink2:#434954;--muted:#6b7280;--hair:#dde0e5; +--accent:#1b5e9b;--accent-soft:#e3eef8;--go:#1e6a4d;--go-bg:#e0f1e8;--warn:#8a5900;--warn-bg:#fbeed3;--stop:#9c2b33;--stop-bg:#f8e2e4; +--sans:"Pretendard","Apple SD Gothic Neo","Malgun Gothic",system-ui,sans-serif;--mono:ui-monospace,SFMono-Regular,Menlo,monospace} +@media (prefers-color-scheme:dark){:root{--paper:#111317;--surface:#1a1d23;--sunk:#22262d;--ink:#eceef2;--ink2:#c2c7d0;--muted:#8e95a2; +--hair:#2d323a;--accent:#79b2e8;--accent-soft:#1a2b3c;--go:#7fd3ab;--go-bg:#15302a;--warn:#e8b45c;--warn-bg:#33291a;--stop:#ee8f96;--stop-bg:#371e21}} +*{box-sizing:border-box}html{scroll-behavior:smooth}body{margin:0;background:var(--paper);color:var(--ink);font:16px/1.7 var(--sans);word-break:keep-all} +a{color:var(--accent)}code{font-family:var(--mono);font-size:.9em;background:var(--sunk);padding:1px 5px;border-radius:4px} +.nav{position:sticky;top:0;z-index:5;background:color-mix(in srgb,var(--paper) 88%,transparent);backdrop-filter:blur(8px);border-bottom:1px solid var(--hair)} +.nav .in{max-width:1080px;margin:0 auto;padding:10px 16px;display:flex;gap:16px;align-items:center;justify-content:space-between;flex-wrap:wrap} +.nav b a{color:var(--ink);text-decoration:none}.nav .links{display:flex;gap:14px;flex-wrap:wrap;font-size:14px}.nav .links a{color:var(--ink2);text-decoration:none} +.nav .links a:hover{color:var(--accent)}.wrap{max-width:1080px;margin:0 auto;padding:0 16px 80px} +.hero{padding:64px 0 28px}.eyebrow{font-size:13px;font-weight:700;letter-spacing:.06em;color:var(--accent)} +h1{font-size:clamp(30px,5vw,48px);line-height:1.18;margin:10px 0 14px;letter-spacing:-.02em} +h2{font-size:clamp(22px,3vw,28px);margin:0 0 6px;letter-spacing:-.01em}h3{font-size:17px;margin:0 0 6px} +.lede{color:var(--ink2);font-size:18px;max-width:62ch;margin:0}.sub{color:var(--ink2);max-width:66ch;margin:0 0 20px} +section{padding:56px 0 8px;scroll-margin-top:56px}.muted{color:var(--muted);font-size:14px} +.cta{display:flex;gap:10px;flex-wrap:wrap;margin:22px 0 0}.btn{display:inline-block;padding:10px 18px;border-radius:999px;font-weight:600;text-decoration:none;font-size:15px} +.btn.primary{background:var(--accent);color:#fff}.btn.ghost{border:1px solid var(--hair);color:var(--ink);background:var(--surface)} +.card{background:var(--surface);border:1px solid var(--hair);border-radius:14px;padding:20px} +.grid{display:grid;gap:14px;grid-template-columns:repeat(auto-fit,minmax(280px,1fr))}.grid3{display:grid;gap:14px;grid-template-columns:repeat(auto-fit,minmax(240px,1fr))} +.icon{width:36px;height:36px;border-radius:10px;background:var(--accent-soft);color:var(--accent);display:grid;place-items:center;font-weight:800;margin-bottom:10px} .badge{display:inline-block;font-size:12.5px;font-weight:600;padding:2px 9px;border-radius:999px;white-space:nowrap} -.go{color:var(--go);background:var(--go-bg)}.warn{color:var(--warn);background:var(--warn-bg)}.stop{color:var(--stop);background:var(--stop-bg)} +.go{color:var(--go);background:var(--go-bg)}.warn{color:var(--warn);background:var(--warn-bg)}.stop{color:var(--stop);background:var(--stop-bg)}.info{color:var(--accent);background:var(--accent-soft)} .gates{display:flex;gap:6px;flex-wrap:wrap;margin:10px 0 4px} -textarea{width:100%;min-height:74px;font:inherit;padding:12px;border-radius:10px;border:1px solid var(--hair);background:var(--surface);color:var(--ink)} -.hit{outline:2px solid var(--accent)}table{width:100%;border-collapse:collapse;font-size:14px} -th,td{text-align:left;padding:8px 10px;border-bottom:1px solid var(--hair);vertical-align:top}th{color:var(--muted);font-weight:600} -.tablewrap{overflow-x:auto;background:var(--surface);border:1px solid var(--hair);border-radius:10px} -.banner{border-radius:10px;padding:12px 16px;margin:14px 0;font-weight:600} -.kpi{display:flex;gap:28px;flex-wrap:wrap;margin:8px 0}.kpi div{min-width:120px}.kpi .v{font-size:26px;font-weight:700} -.figs{display:grid;gap:12px;grid-template-columns:repeat(auto-fit,minmax(300px,1fr))}.figs img{width:100%;background:#fff;border-radius:8px;border:1px solid var(--hair)} -svg text{fill:var(--ink2);font:12px var(--sans)}ul{padding-left:20px}footer{margin-top:60px;color:var(--muted);font-size:13px;border-top:1px solid var(--hair);padding-top:14px} +.ask{margin-top:26px}.ask textarea{width:100%;min-height:76px;font:inherit;font-size:17px;padding:14px 16px;border-radius:14px;border:1px solid var(--hair);background:var(--surface);color:var(--ink)} +.ask textarea:focus{outline:2px solid var(--accent);border-color:transparent}.chips{display:flex;gap:6px;flex-wrap:wrap;margin:10px 0} +.chip{border:1px solid var(--hair);background:var(--surface);color:var(--ink2);border-radius:999px;padding:4px 12px;font:inherit;font-size:13.5px;cursor:pointer} +.chip[aria-pressed=true]{background:var(--ink);color:var(--paper);border-color:var(--ink)} +.hit{outline:2px solid var(--accent);outline-offset:2px}.topic{text-decoration:none;color:inherit;display:block}.topic:hover{border-color:var(--accent)} +.list .topic{display:grid;grid-template-columns:minmax(0,1.2fr) minmax(0,2fr) auto;gap:14px;align-items:center} +.list{display:flex;flex-direction:column;gap:8px}.list .topic p{margin:0} +table{width:100%;border-collapse:collapse;font-size:14px}th,td{text-align:left;padding:9px 10px;border-bottom:1px solid var(--hair);vertical-align:top}th{color:var(--muted);font-weight:600} +.tablewrap{overflow-x:auto;background:var(--surface);border:1px solid var(--hair);border-radius:12px} +.banner{border-radius:12px;padding:12px 16px;margin:14px 0;font-weight:600} +.figs{display:grid;gap:12px;grid-template-columns:repeat(auto-fit,minmax(300px,1fr))}.figs img{width:100%;background:#fff;border-radius:10px;border:1px solid var(--hair)} +svg text{fill:var(--ink2);font:12.5px var(--sans)}svg .strong{fill:var(--ink);font-weight:700}ul,ol{padding-left:20px} +.flow{display:grid;gap:10px;grid-template-columns:repeat(auto-fit,minmax(150px,1fr));counter-reset:s} +.flow .card{padding:14px;position:relative}.flow .n{font:700 12px var(--mono);color:var(--accent)}.flow b{display:block;margin:2px 0 4px}.flow p{margin:0;font-size:14px;color:var(--ink2)} +.flow .who{margin-top:8px;font-size:12.5px;color:var(--muted)}.split{display:grid;gap:14px;grid-template-columns:repeat(auto-fit,minmax(300px,1fr))} +.vs .card h3{display:flex;gap:8px;align-items:center}.stat{font-size:28px;font-weight:800;letter-spacing:-.02em} +footer{margin-top:64px;color:var(--muted);font-size:13px;border-top:1px solid var(--hair);padding-top:16px} +@media (max-width:640px){.list .topic{grid-template-columns:minmax(0,1fr)}.hero{padding-top:40px}} """ def page(title: str, body: str, depth: int = 0) -> str: up = "../" * depth + home = f"{up}index.html" return f""" {e(title)} -
-
정책 효과 분석 플랫폼 -GitHub · 가짜연구소 인과추론팀 × OpenUp
-{body} -
모든 수치와 판정은 레포의 catalog/·cases/에서 GitHub Actions가 자동 생성합니다. -시뮬레이션 데이터로 만든 결과에는 별도 표시가 붙습니다. · {REPO}
-
""" + + + +
{body} +
가짜연구소 인과추론팀 × 오픈업 오픈소스 AI 특화형 트랙3 · 모든 수치와 판정은 레포의 catalog/·cases/에서 +GitHub Actions가 자동 생성합니다(매주 월요일 갱신). 시뮬레이션 데이터로 만든 결과에는 별도 표시가 붙습니다. +· 소스 코드 (MIT)
""" def gate_badges(t) -> str: @@ -203,9 +222,10 @@ def topic_page(t, policies) -> str: by_id = {p.id: p for p in policies} ps = [by_id[i] for i in t.policies] body = [ - f"

{e(t.name)}

{e(t.question)}

", + f"
주제" + f"

{e(t.name)}

{e(t.question)}

", f'
{gate_badges(t)}
', - "

1. 이 주제의 정책 전체

", + "

1. 이 주제의 정책 전체

", f"

수집 방법: {COLLECT_KO[t.collect.method]} — {e(t.collect.note)}

", timeline_svg(t.events), ] @@ -271,14 +291,45 @@ def topic_page(t, policies) -> str: return page(t.name, "\n".join(body), depth=1) +def readiness(t) -> tuple[str, str, str]: + """주제 전체 준비 상태 (필터용 키, 라벨, 색).""" + st = {g.status for g in t.gates.values()} + if "fail" in st: + return "fail", "관문 탈락", "stop" + if "check" in st: + return "check", "확인 필요", "warn" + if "key" in st: + return "key", "데이터 키 대기", "warn" + return "ready", "분석 가능", "go" + + +def has_result(t, policies) -> bool: + by_id = {p.id: p for p in policies} + return any( + by_id[i].sample_case and (ROOT / by_id[i].sample_case / "flow_log.json").exists() + for i in t.policies + ) + + +FLOW_STEPS = [ + ("소셜 신호", "뉴스 제목·SNS 글에서 사람들이 궁금해하는 것을 읽습니다", "누구나"), + ( + "주제 고르기", + "신호는 주제까지만 정합니다. 화제가 된 정책 하나를 고르지 않습니다", + "플랫폼 자동", + ), + ("정책 전부 모으기", "법제처 조례·부처 고시로 '어느 지역이 언제부터'를 모읍니다", "수집 담당"), + ("세 관문", "언제 시작했나 · 누가 받았나 · 무엇으로 재나", "문제 정의 담당"), + ("계획 먼저", "데이터를 보기 전에 분석계획을 커밋합니다(사전 등록)", "문제 정의 담당"), + ( + "효과 추정·판정", + "식별됨 · 조건부 · 식별 불가 — 과장 표현은 자동으로 막습니다", + "추정·리포트 담당", + ), +] + + def index_page(topics, policies) -> str: - cards = [] - for t in topics: - cards.append( - f'' - f"

{e(t.name)}

{e(t.question)}

" - f'
{gate_badges(t)}
' - ) kw = { t.id: list( dict.fromkeys( @@ -288,38 +339,102 @@ def index_page(topics, policies) -> str: ) for t in topics } - steps = [ - ("소셜 신호", "뉴스 제목·SNS 글"), - ("주제", "신호는 주제까지만 정함"), - ("정책 전체 수집", "법제처 조례·고시로 지역×시점"), - ("세 관문", "언제 · 누가 · 무엇을"), - ("사전 등록", "계획을 먼저 커밋"), - ("효과 추정", "판정: 식별됨·조건부·식별 불가"), + cards = [] + for t in topics: + key, lab, cls = readiness(t) + res = has_result(t, policies) + cards.append( + f'

{e(t.name)}

' + f'{lab}' + + (' 분석 결과 있음' if res else "") + + f'

{e(t.question)}

{gate_badges(t)}
' + ) + n_ready = sum(1 for t in topics if has_result(t, policies)) + flow = "".join( + f'
{i + 1:02d}{a}

{b}

{c}
' + for i, (a, b, c) in enumerate(FLOW_STEPS) + ) + examples = [ + "토허제 확대하고 강남 집값 잡혔나요?", + "지역화폐 쓰면 동네 가게 매출이 오르나요?", + "5030 속도 줄이고 사고 줄었나?", + "계절관리제 하면 미세먼지 줄어요?", ] + chips = "".join(f'' for x in examples) body = f""" -

소셜 반응에서 정책 효과까지

-

사람들이 이야기하는 정책이 실제로 효과가 있었는지, 공공데이터로 확인합니다. -화제가 된 정책 하나만 골라 분석하지 않고, 그 주제의 정책을 전부 모아 비교합니다. -화제성으로 사례를 고르면 결과를 보고 사례를 고르는 셈이 되기 때문입니다.

-
{"".join(f'
{i + 1}{a}{b}
' for i, (a, b) in enumerate(steps))}
-

지금 어떤 이야기가 궁금하세요?

- -

글을 입력하면 해당하는 주제를 찾아 표시합니다. (브라우저 안에서만 동작, 서버로 보내지 않음)

-

주제

{"".join(cards)}
+
OPEN SOURCE · 공공데이터 · 인과추론 +

사람들이 묻는 정책,
정말 효과가 있었을까?

+

소셜 반응에서 출발해, 그 주제의 정책을 전부 모으고, 공공데이터로 효과를 추정합니다. +결론을 낼 수 없으면 "식별 불가"라고 말하는 것까지가 이 플랫폼의 일입니다.

+
+
{chips}
+

입력한 글은 이 브라우저 안에서만 쓰입니다. 서버로 보내지 않습니다.

+ +
+ +

왜 이렇게 만드나

화제가 된 정책만 골라 분석하면 결과를 보고 사례를 고르는 셈이 됩니다.

+
+

A 화제성으로 사례를 고르면

+
  • 조용했지만 효과가 컸던 정책이 빠집니다
  • 화제가 되면 신청이 늘어 효과가 부풀려집니다
  • +
  • 시행 전부터 화제였다면 사람들이 미리 움직여 비교가 깨집니다
+

B 화제성은 출발점으로만 쓰면

+
  • 신호는 어느 주제를 볼지만 정합니다
  • 그 주제의 정책을 전부 모아 비교합니다
  • +
  • 신호는 나중에 '미리 반응했나' 점검용으로 다시 씁니다
+ +

여섯 단계 흐름

각 단계는 조원 역할과 7주 일정에 그대로 대응합니다.

+
{flow}
+ +

주제

주제마다 세 관문(언제·누가·무엇을)의 통과 여부를 보여줍니다.

+
+ + + + + +
+
{"".join(cards)}
+ +

무엇을 믿을 수 있나

결과보다 과정을 먼저 공개합니다.

+
+
1

계획을 먼저 커밋

분석계획(plan.yaml)이 git에 커밋되지 않으면 추정 단계가 실행되지 않습니다.

+
2

방법은 규칙이 정함

LLM은 주제·서술을 돕고, 추정 방법은 데이터 모양을 보고 규칙이 고릅니다. 수치는 검증된 라이브러리가 계산합니다.

+
3

작은 표본에 맞는 추론

처치 지역이 10곳 미만이면 무작위화 추론으로 바꿉니다. 시뮬레이션에서 95% 신뢰구간 포함률 95%를 확인했습니다.

+
4

과장 표현 차단

판정이 '식별됨'이 아니면 "입증", "때문에" 같은 표현을 리포트에서 막습니다.

+
5

출처·라이선스 표시

모든 데이터셋에 제공 기관, 공공누리 유형, 수집 시점을 붙입니다.

+
6

한계도 공개

시뮬레이션 결과, 키 대기, 확인 필요 상태를 숨기지 않고 배지로 표시합니다.

+
+ +

참여하기

조별로 주제 하나를 맡아 cases/에 케이스를 추가합니다.

+
+
①

주제 고르기

위 주제 중 하나를 고르거나 catalog/topics.yaml에 새 주제를 제안합니다.

+
②

계획 PR

cases/_template을 복사해 plan.yaml을 쓰고 PR로 사전 등록합니다.

+
③

실행·공개

make flow로 돌리면 이 사이트에 결과가 자동으로 올라옵니다.

+
+ """ return page("정책 효과 분석 플랫폼", body) @@ -332,6 +447,12 @@ def main() -> None: (OUT / "index.html").write_text(index_page(topics, policies), encoding="utf-8") for t in topics: (OUT / "topics" / f"{t.id}.html").write_text(topic_page(t, policies), encoding="utf-8") + sys.path.insert(0, str(Path(__file__).resolve().parent)) + import architecture + + (OUT / "architecture.html").write_text( + page("아키텍처 · 정책 효과 분석 플랫폼", architecture.body(e)), encoding="utf-8" + ) (OUT / ".nojekyll").write_text("") print(f"_site/ 생성: 주제 {len(topics)}개") diff --git a/tests/test_site_build.py b/tests/test_site_build.py index 2c23aa2..614e919 100644 --- a/tests/test_site_build.py +++ b/tests/test_site_build.py @@ -14,7 +14,7 @@ def test_site_builds(tmp_path, monkeypatch): mod.main() out = tmp_path / "_site" topics = mod.load_topics() - assert (out / "index.html").exists() + assert (out / "index.html").exists() and (out / "architecture.html").exists() for t in topics: page = (out / "topics" / f"{t.id}.html").read_text(encoding="utf-8") assert t.name in page