Your Meeting Transcripts Are Training Data: What the EDPB's July 2026 Guidelines Mean for Otter, Fireflies, and Granola Users
Published August 06, 2026
- Otter, Fireflies, and Granola all train on customer meeting content by default on non-Enterprise plans — the opt-out is per-user and easy to miss.
- The EDPB's Guidelines 03/2026 (adopted July 7-8, 2026) confirm GDPR applies in full to AI training data, with consent 'will most probably not' work at scale.
- Once a transcript enters a training pass, the contribution to model weights is durable — deletion does not unwind it.
- The Chamberlain v. Granola complaint (filed July 30, 2026) alleges training-by-default is a deliberate design choice, not an accident.
- On-device transcription eliminates the entire training-corpus surface: no vendor server ever holds the audio.
Quick answer: Yes — Otter, Fireflies, and Granola all train on customer meeting content by default on most plans, with only user-level opt-outs. The EDPB's July 2026 draft guidelines confirm the GDPR applies in full to this training, rule out consent as a workable basis at scale, and require a documented three-part legitimate-interest test that a silent notetaker is structurally unlikely to pass.
If you're using Otter, Fireflies, or Granola on anything other than an Enterprise plan, your meeting transcripts are — right now, by default — being used to train the vendor's AI models. On July 7-8, 2026 in Brussels, the European Data Protection Board adopted Guidelines 03/2026 on web scraping in the context of generative AI — the first pan-EU framework confirming that the GDPR applies in full whenever personal data is used to train AI, with no carve-out. The guidelines were written for web scrapers, but the principles land squarely on the AI notetaker industry: consent "will most probably not" work at scale, legitimate interest requires a documented three-part test, and once data enters model weights, it can't easily be removed. This piece walks through what changed on July 7, what your notetaker actually does with your conversations, and why on-device transcription is the only architecture that avoids the problem entirely.
What the EDPB actually adopted on July 7-8, 2026
The European Data Protection Board is the panel of national data-protection regulators that coordinates GDPR enforcement across the EU. At its Brussels plenary, per YuSMP Group's regulatory summary, it adopted Guidelines 03/2026 on web scraping in the context of generative AI — a 22-page framework confirming that the GDPR applies in full to personal data scraped to train AI models, with no carve-out for AI. Public consultation runs until October 30, 2026, but the enforcement direction is already clear.
Reporting from PPC Land characterized the guidance as blocking AI firms from using consent as an excuse to scrape, setting out a three-condition legitimate-interest test that firms will have to document dataset-by-dataset. Reed Smith's analysis notes the EDPB expressly acknowledged that this kind of processing typically occurs without the data subject's knowledge or consent, creating significant risks to fundamental rights.
The four things the guidelines say that matter for notetakers
Distilled from the YuSMP summary:
- Consent "will most probably not" work. Public visibility of data — or, by analogy, presence on a Zoom call — is not consent to have that content ingested for AI training.
- Legitimate interest needs a documented three-part test (genuine interest, necessity, and a balancing test the data subjects' rights don't override).
- Special-category data is near-prohibited. Incidental collection — the medical detail that comes up in a client call, the biometric voiceprint generated by speaker diarization — must be minimised and deleted on discovery.
- Once trained, data can't easily be deleted from a model. Governance must happen upstream, before ingestion, not after.
That last point is the sharp end for anyone whose voice has ever been on a call recorded by a mainstream cloud notetaker.
What your notetaker actually does with your meetings
The default training behavior varies by vendor and plan, and it's almost never the answer you'd guess.
Otter.ai
Otter's own privacy policy discloses sharing with "data labeling service providers who provide annotation services and use the data we share to create training and evaluation data for Otter's product features." A Cyber Unit review of six major notetaker privacy policies found that these labeling companies are not named in Otter's privacy policy, and standard users have no opt-out for this transfer. The same review confirmed Otter also works with Facebook and other advertising partners for personalized ad targeting.
The Otter lawsuit filings go further. According to a structured analysis of the Brewer complaint, Otter's Privacy Policy admits to training its AI on de-identified recordings and transcripts but only seeks consent from users, not from non-users present in those meetings — effectively outsourcing its legal obligations by asking hosts to "make sure you have the necessary permissions."
Granola
Granola's positioning as a "bot-free" alternative was supposed to sidestep the notetaker consent problem. Then Chamberlain v. Granola was filed on July 30, 2026. According to the PPC Land reporting on the complaint, model training is enabled by default for users on Granola's Free and Business plans, through a setting a user must locate and manually disable. On Business plans, each individual user must opt out separately rather than an administrator doing so once for the organization.
Granola's own security page confirms the training default: "Granola trains on your anonymized data so we can keep making Granola better. You can opt out of this in your Settings." Only Enterprise users have model training turned off by default. A detailed Granola privacy audit published August 3, 2026 confirmed that training on your meetings is on by default for the free tier and for Business, and off by default only for Enterprise.
Fireflies
Fireflies is already a co-defendant in the consolidated Northern District case. Per the Embertype notetaker-lawsuit tracker, Cruz v. Fireflies.AI Corp. was filed on December 18, 2025, raising similar consent-and-recording claims to the Otter lawsuit and consolidated before the same federal judge. Its privacy policy should be read alongside the complaint before any regulated deployment.
Comparison: how training defaults stack up
| Vendor / plan | Audio location | Training on your content? | Opt-out mechanism | Third-party AI providers |
|---|---|---|---|---|
| Otter (Basic / Pro) | Otter cloud | Yes, default | No user-facing training toggle documented; data-labeling transfer has no standard opt-out | Yes, unnamed labeling providers |
| Granola (Free / Business) | Granola cloud (audio deleted post-transcription) | Yes, default | Per-user Settings toggle | Deepgram, AssemblyAI, OpenAI, Anthropic — contractually barred from training |
| Granola Enterprise | Granola cloud | Off by default org-wide | Admin-enforced | Same, contractually barred |
| Fireflies | Fireflies cloud | Yes, subject to plan / policy; in active litigation | Plan-dependent | Yes |
| Basil AI | Your iPhone / Mac only | No — architecturally impossible | N/A — no server holds the audio | None. Apple Speech Recognition, on-device. |
Why "we anonymize before training" isn't a defense
Every major notetaker's response to the training question is some version of "we train on de-identified data." The Brewer complaint's structured critique cites academic research showing that ML models trained on de-identified data remain vulnerable to re-identification attacks, rendering the de-identification claim weak. Voice content is particularly bad here: a voiceprint isn't just personal data, it's biometric identifier data — the very category the EDPB flagged as near-prohibited absent both an Article 6 basis and an Article 9(2) exception.
And the durability problem is real. A detailed retention analysis of Otter's data lifecycle put it bluntly: once a meeting transcript has contributed to a model-training pass under the opt-out default, the de-identified contribution to model weights is durable, and deleting the original recording does not unwind the training contribution.
Article 5 GDPR and the transparency problem
Article 5 of the GDPR requires that personal data be processed "lawfully, fairly and in a transparent manner in relation to the data subject." A meeting participant who never signed up for the notetaker, never saw a disclosure, and had no chance to decline before the transcription started cannot, by any reading of the text, be said to have received transparent processing.
The National Law Review's analysis of the Otter litigation makes a related point about California law that reads across to Europe: the fact that a notetaker does not store audio does not eliminate exposure. The EU regime — Article 5 transparency plus the new EDPB guidance on training-data lawful basis — points in the same direction. Ephemeral processing that still results in personal data entering a training corpus is still processing that needs a lawful basis.
The EU AI Act layer arriving August 2, 2026
The EDPB guidance doesn't sit alone. HR Executive's coverage of the Littler analysis flags that beginning in August 2026, the EU AI Act introduces a separate layer of obligation, and AI systems used for worker monitoring and management may be classified as high-risk — a category that could encompass tools offering sentiment analytics or productivity scoring alongside transcription.
The same reporting notes that in co-determination countries such as Germany and France, deploying an AI notetaker may also require works council consultation before rollout, a requirement with no U.S. equivalent that multinational HR teams frequently overlook. For further coverage of how these frameworks interact, see our analysis of Article 50 transparency obligations for AI notetakers.
The litigation backdrop: three consolidated cases
The EDPB guidelines land in the middle of an already-hot U.S. docket. Per the Embertype lawsuit tracker, since August 2025 three major federal privacy lawsuits have been filed against AI notetaker companies. Recording Law's docket summary notes that Judge Eumi K. Lee consolidated the Otter suits into In re Otter.AI Privacy Litigation, No. 5:25-cv-06911-EKL, with a consolidated complaint filed December 5, 2025 alleging Otter obtained consent, at most, from the meeting host who added the assistant — not from the other participants whose voices it recorded and transcribed.
The complaint further alleges that Otter used recorded content to train and improve its artificial intelligence models. The case docket summary confirms Otter moved to dismiss, the motion-to-dismiss hearing was held on May 20, 2026, and Judge Lee took the matter under submission. California's CIPA allows statutory damages of $5,000 per violation, and the federal Wiretap Act provides its own statutory damages.
Chamberlain v. Granola, filed July 30, 2026 in the same district, extends the theory to bot-free notetakers. Per the PPC Land complaint summary, the filing extends California's wiretapping and privacy statutes to an AI product whose core function is recording conversations and repurposing them for model training. For a deeper look at that case, see our Chamberlain v. Granola breakdown.
What professionals should do this week
If you handle any content that would embarrass a client, disclose a trade secret, or make discovery counsel nervous — sales calls, therapy sessions, investigator interviews, deposition prep, board discussions — you need to know the training posture of every AI notetaker on every device on your calls. That means:
- Audit your seats. For each Otter / Fireflies / Granola account in your org, verify whether the training toggle is off. On non-Enterprise plans, assume it's on until you check.
- Push to Enterprise or off-cloud. Enterprise plans at least give admins one place to enforce a training default. On-device transcription eliminates the surface entirely.
- Document your legitimate-interest test if you're an EU controller relying on any vendor that trains. The EDPB expects that documentation to exist before ingestion, not after a complaint.
- Refresh calendar-invite language to disclose AI-assisted transcription. This is a bare minimum for CIPA/BIPA/GDPR transparency and does not, by itself, resolve the training-consent problem for non-users on the call.
For a broader buyer's framework, see our companion piece on what "compliant AI meeting notes" actually means.
How Basil AI solves this
Basil AI was built for exactly this problem. Every transcription runs on-device using Apple's Speech Recognition framework and the Neural Engine — see Apple's Speech framework documentation and Apple's privacy overview. Concretely:
- No vendor server holds the audio. There is no cloud copy of your meeting to include in a training corpus, subpoena, or breach notification.
- No training on your content — architecturally. Basil doesn't have a training pipeline that ingests customer meetings, because it doesn't have customer meetings on its infrastructure.
- No third-party AI providers with access. Not Deepgram, not AssemblyAI, not OpenAI, not Anthropic. Transcription runs on your device; summarization uses Apple's on-device foundation models where available.
- No opt-out toggle to hunt for. The training question doesn't arise. That is the whole point.
Compliance is your determination as the controller. What on-device architecture gives you is a cleaner factual record when you make it: no vendor server holding the recording, no processor to add to your Article 30 register for the transcription step, and no training-corpus question to answer on your DPIA. For the technical detail, see our deep dive on on-device transcription on iOS 26.
The bottom line
The EDPB didn't need to name AI notetakers to change their exposure. Guidelines 03/2026 confirms what has always been true under the GDPR — that scraping personal data at scale to train AI needs a lawful basis, that consent doesn't work at that scale, and that legitimate interest requires documentation your average SaaS notetaker doesn't have. Combine that with the pending U.S. wiretap and BIPA cases, the training-by-default posture that reporting keeps uncovering on the top three vendors, and the durability of model-weight contributions, and the direction of travel is clear. The safest meeting transcript is the one that never leaves your device.
Keep your meetings out of anyone's training data
Basil AI transcribes 100% on-device. No cloud upload, no training pipeline, no third-party AI providers. Your voice never leaves your iPhone or Mac.
Frequently Asked Questions
Do AI notetakers train on my meeting transcripts by default?
On most consumer and mid-tier plans, yes. Granola's own filings and reporting confirm training is enabled by default on Free and Business plans, requiring each user to locate and toggle it off. Otter's privacy policy authorizes training on de-identified recordings and transcripts. Enterprise plans typically flip the default, but only Enterprise — everyone else contributes to model training until they opt out.
What did the EDPB actually say on July 7-8, 2026?
The European Data Protection Board adopted Guidelines 03/2026 on web scraping for generative AI. It confirms the GDPR applies whenever personal data is used to train AI models, states that consent 'will most probably not' work at scale, and imposes a documented three-part legitimate-interest test (interest, necessity, balancing). Public consultation runs until October 30, 2026.
Do the EDPB guidelines apply to meeting transcripts, not just web scraping?
The guidelines target web scraping directly, but the underlying GDPR principles — lawful basis, transparency, data minimization, and the difficulty of purging data from trained model weights — apply identically to any large-scale training corpus built from personal data, including meeting transcripts captured from non-consenting participants. Regulators will read across.
Can I get my voice out of a model once it's been trained?
Practically, no. As legal analysts covering Otter noted, once a transcript contributes to a training pass, the de-identified contribution to model weights is durable. Deleting the original recording does not unwind the training contribution. The EDPB guidance explicitly warns that governance must be upstream because trained data can't easily be deleted from a model.
How is Basil AI different?
Basil AI transcribes 100% on-device using Apple's Speech Recognition and Neural Engine. Audio never leaves your iPhone or Mac, so there is no vendor server to hold recordings, no training corpus built from your meetings, no third-party AI providers with access, and no opt-out toggle to hunt for — because there is nothing to opt out of.
Does an Enterprise plan solve the training problem?
It reduces it, but doesn't eliminate the underlying architecture. Enterprise plans at Granola turn off model training org-wide by default, and Otter's enterprise contracts can carry stronger guarantees. But audio and transcripts still transit the vendor's cloud, are still subject to subpoena, backup retention, and legal-hold pauses, and still depend on the vendor's contractual promises rather than on-device architecture.