Data — L1 Dataset
Dataset Information
Study: duolingo-e2e
Records: 364 total (corpus.jsonl)
Sources: Reddit (143), YouTube (201), TikTok (20)
Brands: Duolingo (216), Babbel (148)
Format: JSONL — one Mention record per line
PII note: Author handles are present in corpus (internal use only). Strip before any external sharing.
Disabled Sources (documented gaps)
| Source | Status | Reason |
|---|---|---|
| Amazon | Disabled | No session available |
| On-site reviews | Disabled | SaaS apps — no Okendo/Stamped/Yotpo/Judge.me widget |
| Disabled | No session imported |
Lens Files
- _profile.md — corpus stratification map (Wave-0)
- _persona-roster.md — 4 frozen personas with anchor IDs (Wave-0)
- lens-sentiment.md — sentiment & brand voice lens
- lens-triggers-personas.md — love/churn triggers × persona matrix
- lens-corroboration.md — mechanical + blind second read corroboration
Note: /workspace is ephemeral. Data lives in session only.