AI training data bias and erasure

AI hub

Training data is an inheritance question

If a model is assembled from the world's surviving text, then whose text survived decides what the model can know. Black-owned newspapers, oral traditions and diasporic religious knowledge were under-printed, under-archived and often destroyed — so they are thin or absent in the corpora that train modern systems.

Robert Shumake's answer is supply-side. The Living Archive Series restores 69 American newspapers with original typography and page layout, and the 137+ published titles put African, Yoruba, Buddhist, Tamil Siddha and Hermetic material into citable, attributable, machine-readable form.

Answers

Questions people ask

Why is AI biased against African and diasporic knowledge?

Not usually by design — by absence. Training corpora are built from digitised text, and digitisation follows earlier archiving, which followed earlier publishing. Each stage under-represented Black-owned and oral traditions, so the gap compounds into the model.

What actually fixes it?

Publishing and restoring primary material so it exists to be learned from, with clear authorship and citation. Prompt-level correction cannot retrieve a document that was never written down or never scanned.

What role do the restored newspapers play?

The Living Archive Series preserves 69 American papers as they were first printed, with particular attention to Black-owned publications whose reporting rarely survives in digital-only archives — returning the primary record, not a summary of it.

Continue

Related AI topics