📅 August 2, 2026 · ⏱️ 11 min read · By Basil AI Team
Your Podcast, Their Training Data: What the May 2026 BIPA Voiceprint Lawsuits Against Big Tech Mean for Anyone Whose Voice Is Online
Published August 02, 2026
- Nine coordinated BIPA class actions were filed May 11-13, 2026 in the Northern District of Illinois against Adobe, Alphabet/Google, Amazon, Apple, ElevenLabs, Meta, Microsoft, NVIDIA, and Samsung over voiceprints extracted for AI voice-model training.
- BIPA statutory damages are $1,000 per negligent violation and $5,000 per intentional violation — with no requirement to prove actual harm.
- The novel legal theory: even publicly available audio (podcasts, newscasts, audiobooks) is protected because BIPA regulates the biometric identifier, not the recording.
- The same voiceprint-collection theory drives separate suits against Fireflies.ai, Otter.ai, and by extension any cloud transcription tool that performs speaker diarization.
- On-device transcription — where audio and speaker-differentiation math never leave the device — is the only architecture that sidesteps the collection-based BIPA theory entirely.
Quick answer: No — not in Illinois. In May 2026, nine coordinated BIPA class actions were filed in the Northern District of Illinois against Adobe, Alphabet/Google, Amazon, Apple, ElevenLabs, Meta, Microsoft, NVIDIA, and Samsung, alleging they extracted voiceprints from publicly available recordings to train commercial AI voice models without written notice or consent. BIPA requires consent even when audio is public.
If you have ever released a podcast, narrated an audiobook, appeared on a newscast, or posted a voice memo online, your voice may already be inside a commercial AI model — and as of May 2026, that is now a federal class-action question. Between May 11 and May 13, 2026, plaintiffs filed nine coordinated class actions in the U.S. District Court for the Northern District of Illinois against Adobe, Alphabet/Google, Amazon, Apple, ElevenLabs, Meta, Microsoft, NVIDIA, and Samsung, alleging each company built commercial AI voice systems on voiceprints extracted from publicly available recordings without the written consent, notice, or retention schedule required by the Illinois Biometric Information Privacy Act (BIPA).
The novel legal theory — and the reason every general counsel, podcaster, and product team should read the complaints — is that public availability of an audio file is legally irrelevant. BIPA regulates the biometric identifier (the mathematical voiceprint), not the recording it was derived from. This piece walks through what was filed, what BIPA actually requires, what the statutory damages look like, how the theory extends to everyday AI meeting tools like Otter and Fireflies, and why on-device transcription is the only architecture that sidesteps the collection-based theory entirely.
What was actually filed in May 2026
According to an American Bar Association analysis of the filings, a group of seven award-winning broadcast journalists, podcasters, audiobook narrators, and voice actors filed a coordinated set of nine class actions on behalf of themselves and other individuals whose voice recordings were produced or recorded in Illinois. The complaints allege that the defendants extracted plaintiffs' voiceprints from publicly available audio recordings without consent, then used those voiceprints to train commercial AI voice models.
Two of the lead case numbers give a sense of the docket: Amer v. Eleven Labs Inc., No. 1:26-cv-05437 (N.D. Ill. May 11, 2026), and Marin v. Alphabet, Inc., No. 1:26-cv-05436 (N.D. Ill.). Law360's coverage describes the plaintiffs as journalists and voice actors accusing Google, Meta, Microsoft, NVIDIA, and ElevenLabs of wrongly using their voices to train AI models. Biometric Update confirmed the parallel actions against Amazon, Apple, Adobe, and Samsung.
The named plaintiffs
Per Capitol News Illinois, plaintiffs include journalist Robin Amer, audiobook narrators and voice actors Lindsay Dorcus and Victoria Nassif, and podcasters Yohance Lacour and Alison Flowers — all Illinois residents. Their attorney, Ross Kimbarovsky of Loevy + Loevy, told reporters the companies had "built a billion-dollar industry on stolen voices," calling the alleged conduct one of the largest violations of biometric privacy ever committed.
Why "public audio" doesn't save the defendants
The instinct — and probably the first defense — is that a podcast someone voluntarily released to the world is fair game for scraping. BIPA does not agree. The statute regulates the collection, use, storage, and destruction of biometric identifiers, which the Illinois legislature defined to expressly include voiceprints. The recording is not the identifier. The mathematical model derived from the recording — the pitch, cadence, tone, and vocal-tract characteristics that make your voice recognizably yours — is the identifier.
Mason LLP's analysis puts it concisely: a voiceprint is a mathematical model of the unique characteristics of your voice — pitch, cadence, tone, and vocal-tract shape — and much like a fingerprint, it can identify you with a high degree of certainty. Under 740 ILCS 14, companies collecting voiceprints must provide written notice, obtain informed written consent before collection, publish a retention schedule, and destroy the data on schedule.
The competitive-harm angle
The May 2026 complaints add an unusually pointed commercial-injury allegation. Biometric Update reports that in Google's case, plaintiffs claim the technology built on their biometric voiceprints is now used to compete with their professions — Google Text-to-Speech is used by audiobook publishers as an alternative to human narration, while NotebookLM Audio Overviews can be used to generate podcasts, directly competing with investigative audio journalism and narration work.
What BIPA actually requires — the four hooks
Section 15 of BIPA is short but load-bearing. There are four independent obligations, and the complaints allege the defendants violated all of them:
- 15(a) — Retention schedule: Maintain a publicly available written policy establishing a retention schedule and guidelines for permanently destroying biometric identifiers when the initial purpose is satisfied or within 3 years of last interaction, whichever comes first.
- 15(b) — Informed written consent: Before collection, inform the subject in writing that biometric data is being collected, state the specific purpose and length of collection/storage, and receive a written release.
- 15(c) — No profit from biometrics: Do not sell, lease, trade, or otherwise profit from a person's biometric identifier.
- 15(d) — Disclosure limits: Do not disclose or disseminate a person's biometric data without consent or a narrow statutory exception.
Missing any one is an independent violation. The Fireflies.ai complaint discussed by Jackson Lewis's Workplace Privacy Report alleged three of the four in a single filing — no publicly available retention schedule, no written notice of purpose or duration, and no written release from non-account-holders present in recorded meetings.
The damages math is why this matters
BIPA is one of the few state privacy statutes with a private right of action and per-violation statutory damages that do not require proof of actual harm. As tl;dv's litigation tracker summarizes, damages sought in the parallel Fireflies action are $1,000 per negligent violation and $5,000 per reckless or intentional violation, plus attorneys' fees and injunctive relief.
For a voice-model training corpus that may contain tens or hundreds of thousands of Illinois speakers, the aggregate exposure is enormous. History supports the point. Loevy + Loevy's announcement notes that Meta previously reached a $650 million settlement in a 2021 BIPA claim over collection and storage of biometric information, and Google settled with the State of Texas for $1.375 billion in November 2025 to resolve claims arising from unlawful collection of voiceprint and facial geometry through Google Photos and Google Assistant.
Cloud AI vs. on-device: which architecture creates BIPA exposure?
The important question for anyone deploying transcription or voice AI is architectural, not policy. Under a collection-based statute, the risk begins the moment a biometric identifier is generated by an entity's systems. Where the compute happens determines who is a "collector."
| Dimension | Cloud transcription (Otter, Fireflies, Zoom AI Companion, cloud voice models) | On-device transcription (Basil AI) |
|---|---|---|
| Where audio is processed | Vendor's servers | Your iPhone/Mac Neural Engine |
| Voiceprint generated by | Vendor (a BIPA "private entity") | On-device model on your hardware |
| Third-party retention | Yes — subject to vendor policy and subpoena | None |
| Training on your voice | Alleged in In re Otter.AI Privacy Litigation and Cruz v. Fireflies.AI | No vendor training data leaves the device |
| BIPA Section 15(b) consent burden | Vendor (and often the deploying organization) must obtain written consent from every speaker, including non-users | No third-party collector; consent theory doesn't reach a device the speaker controls |
| Subpoena/discovery surface | Vendor's DPAs and cloud copies are discoverable | No vendor copy exists to produce |
This is not a compliance guarantee — BIPA compliance is always the deploying organization's determination — but architecturally, if no vendor server ever generates or stores the voiceprint, the collection-based theory driving Cruz, Fricker, and the May 2026 wave has nothing to attach to.
How the theory extends to your everyday Zoom call
The May 2026 lawsuits are about training data, but the same voiceprint theory has already been applied to real-time meeting tools. As we covered in our earlier analysis of the BIPA lawsuit wave hitting AI meeting bots, plaintiffs in Cruz v. Fireflies.AI Corp., No. 3:25-cv-03399 (C.D. Ill.), and Fricker v. Fireflies.AI Corp., No. 1:26-cv-02675 (N.D. Ill.), argue that Fireflies' "Speaker Recognition" feature generates voiceprints in exactly the way BIPA regulates.
Lewis Rice's April 2026 analysis flags the doctrinal expansion: even where a company does not market its product as "biometric," functionality that distinguishes speakers based on vocal characteristics may qualify as voiceprint collection under Illinois law. The consolidated In re Otter.AI Privacy Litigation, 5:25-cv-06911 (N.D. Cal.), applies the same theory to Otter's bot — as our breakdown of the Otter litigation's May 2026 hearing details, on-device transcription sidesteps every theory in the case because there is no interception, cloud retention, voiceprint, or training data.
What Fireflies, Zoom, and Otter say in their own policies
Reading the vendor policies is instructive. Fireflies' privacy policy describes broad collection of meeting content and voice data associated with user accounts. Otter's privacy policy grants Otter broad rights to process user content, which is why the consolidated federal complaint includes a training-data theory. Zoom's privacy policy is the most restrictive on training use but still relies on cloud processing of audio, and admin controls — not the speaker — determine whether AI Companion is on. If you're a non-account-holder participating in someone else's meeting, none of these policies were shown to you before your voice was captured.
What GDPR and other regimes add on top
Illinois is the sharpest sword in the U.S., but it is not the only one. In the EU, biometric data is a special category under Article 9 of the GDPR, with explicit consent generally required for processing. Article 5's data-minimisation principle further disfavors slurping voice recordings into a training pipeline that keeps them indefinitely. In healthcare, the same audio slurped by a stray bot can trigger a HIPAA breach notification — a scenario documented in the Ontario hospital case we covered in our shadow-AI incident analysis.
How Basil AI solves this
Basil AI runs transcription and speaker separation entirely on-device using Apple's Speech Recognition framework and the Neural Engine on your iPhone or Mac. That has three consequences directly relevant to the May 2026 BIPA theory:
- No third-party collector. The voice math never reaches a Basil server. BIPA regulates "private entities" that collect biometric identifiers; a tool whose model runs on hardware the speaker or their counterparty owns removes the vendor from the collection chain.
- No training corpus. Basil does not, and cannot, ingest your meetings into a training pipeline. The training-data theory in In re Otter.AI Privacy Litigation has no analog when the model doesn't leave your device.
- No subpoena target and no data-broker resale surface. If there is no cloud copy, there is nothing to produce in discovery and nothing to sell.
For a deeper technical walk-through, see our companion piece on how Basil processes audio locally. What Basil's architecture does not solve — and no product can solve — is your independent obligation under state consent statutes to tell counterparties you're recording. Illinois, California, Florida, Massachusetts, and other all-party-consent states still require the human on the call to obtain consent to record, per the Goodwin Procter analysis of AI transcription tools under scrutiny.
Action checklist by role
For voice professionals (journalists, narrators, podcasters, voice actors)
- If your recordings were made or produced in Illinois, contact BIPA class counsel — the nine May 2026 filings expressly reference "other individuals whose voice recordings were produced or recorded in Illinois."
- Add explicit no-AI-training language to distribution agreements and terms on your own sites.
- Do not assume Creative Commons or open podcast feeds authorize biometric extraction. They don't.
For general counsel and CCOs
- Inventory every SaaS tool with a "speaker recognition," "voice ID," or diarization feature. Speaker labels in a transcript are the tell.
- Verify your vendor DPAs include voiceprint-specific representations — most do not.
- For Illinois employees or Illinois-hosted meetings, treat any cloud diarization vendor as a BIPA "private entity" that requires 15(b) written notice and consent.
- Where possible, prefer on-device transcription for confidential internal meetings — it removes the vendor from the collection chain.
For product and engineering teams building voice AI
- Publish a retention schedule that satisfies 15(a) before shipping any diarization or voice-ID feature.
- If you scrape public audio for training, treat every clip whose speaker was in Illinois at the time of recording as requiring separate 15(b) consent.
- Consider whether on-device inference — Apple Neural Engine, Core ML, or WebGPU-based Whisper — meets the product requirement without generating a BIPA collection event on your infrastructure.
The bottom line
The May 2026 filings will take 12 to 24 months to resolve, and BIPA jurisprudence around AI training is genuinely unsettled. But the direction of travel is not. Illinois has already produced multibillion-dollar biometric settlements. Class counsel has a proven playbook. And every AI voice product built on scraped audio — or every meeting tool that runs speaker diarization in the cloud — now has a live theory attached to it. Architecture, not policy, is what determines exposure. If your voiceprint never leaves your device, no one has to sue over what happened to it.
Record meetings without ever creating a cloud voiceprint
Basil AI is the fully on-device AI note-taker for iPhone and Mac. Real-time transcription, speaker separation, 8-hour recording — all on Apple Neural Engine, none of it in the cloud.
Frequently Asked Questions
Does Illinois BIPA protect voice recordings that are already public?
Yes. BIPA (740 ILCS 14) protects the biometric identifier — the mathematical voiceprint — not the recording itself. The May 2026 complaints argue that extracting a voiceprint from a public podcast, audiobook, or newscast still requires written notice, informed consent, and a published retention schedule under Section 15 of BIPA, regardless of whether the source audio was freely available online.
Which companies were sued in the May 2026 BIPA voice-training cases?
Nine coordinated class actions were filed in the Northern District of Illinois between May 11-13, 2026 against Adobe, Alphabet/Google, Amazon, Apple, ElevenLabs, Meta, Microsoft, NVIDIA, and Samsung. Plaintiffs include Illinois journalists, podcasters, audiobook narrators, and voice actors. Lead cases include Amer v. Eleven Labs Inc. (1:26-cv-05437) and Marin v. Alphabet, Inc. (1:26-cv-05436).
How much money is at stake per BIPA voiceprint violation?
BIPA sets statutory damages at $1,000 per negligent violation and $5,000 per intentional or reckless violation, plus attorneys' fees and injunctive relief. Plaintiffs do not need to prove actual harm. With voice-model training corpora containing millions of speakers, aggregate exposure can be enormous — Meta previously paid $650 million to settle a BIPA facial-geometry claim, and Google settled Texas voiceprint claims for $1.375 billion.
Can I sue if I don't live in Illinois but my voice was used?
Generally no, unless the voiceprint collection occurred in Illinois or your recording was made in Illinois. BIPA is Illinois state law and applies to biometric collection tied to the state. Texas and Washington have their own biometric statutes but with weaker private rights of action. Federal biometric privacy legislation has been proposed repeatedly but not enacted, leaving most Americans without a direct remedy.
Do AI meeting transcription tools create voiceprints too?
Yes. Any tool that performs speaker diarization or 'speaker recognition' — including Fireflies.ai, Otter.ai, Zoom AI Companion, and Google Meet's Gemini features — necessarily generates voice-derived identifiers to distinguish who said what. Separate BIPA suits (Cruz v. Fireflies.AI, In re Otter.AI Privacy Litigation) target this same category of collection during virtual meetings, not just public audio scraping.
How does on-device transcription avoid voiceprint liability?
On-device processing means the audio and any speaker-differentiation math never leave your iPhone or Mac. There is no vendor server collecting, storing, or training on your voice. Because BIPA regulates collection and storage by an entity, a tool that never transmits biometric data to a third party sits outside the collection-based theory driving the current wave of cases. Basil AI runs entirely on Apple's Neural Engine.