Last updated: 21 July 2026
What the corpus is
The archive is a structured collection of audio, video, and written testimony about the lived experience of the AI era, together with transcripts, contribution dates, and the consent record attached to each submission. It is built as a research-grade historical archive — not a content feed, not an advertising asset, and not a training corpus by default.
The consent basis
- Every submission is governed by the exact consent text its contributor accepted, which is versioned and snapshotted alongside the submission. We can always show what was agreed to, and when.
- Publication is two-gated: explicit public-display consent from the contributor and review by a human moderator. Neither alone is sufficient. Everything else stays in the private vault.
- Contributors choose how they appear — real name, alias, or anonymous — and can permanently erase their submission at any time via self-serve erasure. Erasure is a true purge: database row and media files, public and private.
What we do — and do not do — with the corpus
- Moderation: each submission is transcribed and given a first-pass automated safety screen before human review. That screening is the only automated processing applied to private submissions.
- Aggregate analysis: the public Pulse is computed only over approved submissions (those whose contributors consented to archive inclusion), and only as aggregate themes — never as individual attribution of private material.
- No sale of personal data. No advertising use.
- No AI-training use of private submissions. Private vault material is processed for moderation and preservation only.
Research access and future dataset releases
The long-term purpose of the archive is scholarship: psychology, sociology, history, ethics, public policy, and human–computer interaction. Any future research access or dataset release will honor, for each submission, the specific consent version its contributor accepted — and erasure requests propagate: a contributor who erases their voice is removed from the operating corpus and from any dataset release we control thereafter. If a partnership (for example, an institutional steward for the archive) changes how the corpus is governed, this page will say so before the change takes effect.
Continuity and durability
- The corpus is exportable in full (submissions, transcripts, consent records, moderation trail) at any moment — the archive is not held hostage by its own infrastructure.
- Off-platform backups of the database and media are maintained so that a platform failure cannot erase the record.
- Operational failures that touch consent promises (a failed erasure, a stalled moderation pipeline) alert a human operator directly.
Questions
Write to privacy@humanvoiceproject.com. This page is version-controlled; material changes are dated at the top.