Contents
  1. Purpose and accountable data use
  2. Information inventories across processing boundaries
  3. Rights, restrictions and decision authority
  4. Transparency, consent and individual requests
  5. Minimization and remaining identification risk
  6. Authorization at access and disclosure boundaries
  7. Approval and authoritative record changes
  8. Lineage and proportionate audit evidence
  9. Vendor and model-processing obligations
  10. Retention schedules and restricted preservation
  11. Lifecycle fulfillment across derivatives
  12. Assurance and operating changes
  13. Check understanding
  14. Open questions
  15. Selected talks
  16. References
  17. Talk library
← All topics

Privacy and Data Governance

An AI assistant can read a support case, draft a reply, update a customer record and reuse the conversation for training. Each operation changes the information’s purpose, recipients or consequences. Privacy and data governance establish who may authorize those changes, the boundaries that software must enforce, and the evidence needed to confirm fulfillment across copies, vendors and models.

Purpose and accountable data use

Data governance assigns decision authority, handling rules and evidence obligations across an information lifecycle. Privacy governance addresses effects on people; the broader discipline also covers licensed material, confidential business information and operational records.

Purpose limitation constrains collection and subsequent use to specified purposes under applicable rules. Secondary use introduces another purpose.

A support assistant’s proposed operating agreement separates decisions that share infrastructure.
Operation and benefitInformation and recipientsDecision authority and boundary
Read a case to resolve itAssigned customer’s case; support workerService owner: permit only the requester’s authorized case access.
Draft a customer replyRelevant case facts; approved processor and customerService owner: exclude internal contact details from customer-facing context.
Change customer contact detailsProposed values; authoritative record serviceRecord owner: require separate write authorization.
Train a reusable modelConversation examples; training operatorAccountable purpose owner: require a separate assessment before reuse.
Evaluate support employeesStaff-linked conversations; managementExclude from case-resolution approval; assess consequences separately.

Necessity connects each field to a benefit. Proportionality weighs that benefit against effects on people and less intrusive alternatives. A decision can narrow the task or refuse it rather than approve every technically possible use.

  • Protection does not grant permissionAccurate, encrypted documents can still carry a license that excludes the intended reuse. AI Security addresses protection; permission requires its own assessment.
  • Permission does not establish fitnessAuthorized conversation examples can contain incorrect model-generated answers. Data Quality and Curation addresses whether those records are suitable for training.

Information inventories across processing boundaries

Under the EU GDPR, personal data relates to an identified or identifiable person—the data subject. Identification can be indirect.

The inventory must include inferred attributes, not just submitted fields. Ordinary writing can support predictions about someone’s location or circumstances, creating sensitive information without an explicit identifier.

One case, several governed artifacts

Example

Branches create separate handling obligations.

All stores shown belong to the application except provider inference. The training branch exists only with separate reuse approval. Arrows represent information movement or transformation, not automatic permission.
Read the diagram as text
  • Case record.
  • Retrieval index.
  • Model input.
  • Provider inference. External organization.
  • Generated reply.
  • Response cache.
  • Audit events.
  • Training examples.
  • Case recordRetrieval index: Derive representations.
  • Retrieval indexModel input: Select context.
  • Model inputProvider inference: Approved external disclosure.
  • Provider inferenceGenerated reply: Return generation.
  • Generated replyResponse cache: Retain response.
  • Model inputAudit events: Record selected metadata.
  • Model inputTraining examples: If reuse approved: input.
  • Generated replyTraining examples: If reuse approved: target.
  • Sources and derived profilesInventory case records, uploads and summaries separately. A profile derived from purchases is a new interpretation of those records, not the original observations.
  • Retrieval representationsEmbeddings are numerical representations used for tasks such as similarity search. Inventory them alongside their source text; numerical representation does not establish concealment.
  • Model interactions and learningInference applies fitted model behavior; training changes fitted state using examples. Prompts, outputs and feedback may later become training records, introducing another use. Machine Learning Fundamentals explains the distinction.
  • Operational copiesTelemetry records system behavior and can contain sensitive attributes. Persistent checkpoints and application stores create further copies with lifetimes independent of a model request.

For each artifact, record its owner, purpose, location, recipients and downstream copies. Inventory relationships as well as storage systems: a protected source can feed a differently governed destination.

Rights, restrictions and decision authority

A lawful basis is an applicable legal ground for processing personal data; consent is not universally required.

Public availability does not settle that assessment. Under the EDPB’s AI-model opinion, reliance on legitimate interests requires an interest, necessity and balancing against affected people’s rights. Development and deployment can have different purposes.

Role or authorityMeaningBoundary
Data controllerUnder EU guidance, determines purposes and essential means.Actual activities determine the role; ordinarily an organization, not its employee.
Data processorA separate entity processing on the controller’s behalf.A contractual label cannot override actual activities.
Internal steward or system ownerMaintains agreed definitions, ownership and operating responsibilities.An internal assignment does not determine the organization’s legal processing role.
Access implementerConfigures credentials and resource permissions.Technical capability does not authorize every use.

Copyright is an independent restriction. The US Copyright Office’s May 2025 report explains that downloading and preparing training material can implicate reproduction rights. Licensing may authorize specified uses; fair use depends on circumstances. The report is institutional analysis, not a binding judgment granting permission for a particular corpus.

Under the EU trade-secrets directive, a trade secret must be secret, commercially valuable because of secrecy, and reasonably protected. Unauthorized copying, disclosure or use contrary to confidentiality duties can be unlawful. Statutory exceptions apply, including specified public-interest disclosures. Other confidential information may remain contractually restricted without satisfying that definition. Internal approval cannot override applicable restrictions.

Qualified review needs a concrete processing description.

  • Review recordName purposes, relevant jurisdictions, affected people, data categories, rights holders, operators and intended recipients; assign responsibility for unresolved restrictions.

Transparency, consent and individual requests

A privacy notice explains handling; access rights concern data and handling information. Under GDPR, additional identity evidence must be necessary to resolve reasonable doubts, not routinely demanded.

When relying on consent, retain the purpose, scope, notice version, when and how agreement was obtained, and withdrawal history. The EDPB consent guidelines require demonstrable validity; a checkbox alone is insufficient. Choice must be meaningful, and withdrawal must be as easy as consent without detriment.

RequestOperational consequence
AccessProvide the applicable data and handling information, respecting others’ rights.
RectificationInvestigate disputed accuracy and explain the outcome.
ErasureAddress live copies and backups when the request is valid and no exemption applies.
RestrictionQualifying cases limit processing; storage may continue.
ObjectionConditions depend on the purpose; direct-marketing use must stop.

Withdrawal does not retrospectively invalidate previously lawful processing. Another independently established purpose may continue under its own ground; an organization cannot invent a replacement ground to evade withdrawal. Route the change by purpose, identify affected processing, and retain evidence of what stopped.

For an inferred customer attribute, determine what the record asserts. The ICO’s rectification guidance distinguishes factual errors from clearly recorded opinions. Assess the person’s evidence and consequences of use, restrict processing while checking where appropriate, and explain decisions and challenge routes. Calling every prediction an opinion does not resolve accuracy duties.

  • Accountable handlingAssign an owner who can investigate, authorize correction and communicate the disposition. Context matters: clinical documentation and professional review involve responsibilities that cannot be replaced by a universal consent rule.

Minimization and remaining identification risk

Minimization limits collection, use, exposure and duration to the approved purpose. Omission prevents a copy from being created; later filtering must locate and transform information already present.

For a delivery-status reply, different reductions preserve different capabilities.
ReductionRetained utilityRemaining limitation
Select one authorized caseCase-specific assistanceOther fields within that case can remain sensitive.
Omit employee contact fieldsCustomer-facing status explanationFree text can still disclose personal details.
Aggregate operational countsWorkload monitoringLoses case detail; aggregation alone does not establish anonymity.
Replace names with stable substitutesLink related recordsThe preserved association can still identify someone.

Redaction removes or masks sensitive content. Detection can return sensitive spans and categories separately from replacement policy, allowing removal, category labels or stable substitutes without conflating extraction with privacy validation.

Pseudonymization separates identifying information under safeguards while retaining possible association. Anonymity requires evidence about identifiability in context, not merely transformed appearance.

  • LinkageThe Netflix-history research showed that auxiliary knowledge of a few ratings and approximate dates could link released records to people, exposing additional ratings. Removing names did not remove distinctive combinations.
  • ReconstructionEmbedding reconstruction research recovered text, including names in clinical examples, under tested encoder-access conditions. This establishes a concrete risk, not a recovery rate for arbitrary indexes.
  • InferenceA study illustrates inferring Melbourne from commuting language describing a characteristic turning maneuver. The mechanism uses contextual clues, not an explicit address. Such inference creates a claim about someone; it does not establish that the claim is correct.

Hashing predictable identifiers permits candidate guessing. Apply purpose-driven capture before export, as explained in Observability; collecting fewer records does not make retained content harmless.

Generated derivatives require their own anonymity assessment. Local execution changes who operates infrastructure, but leaves the application responsible for classification, permissions and appropriate use.

Authorization at access and disclosure boundaries

Authentication establishes identity; authorization determines permitted operations on particular resources. A service account’s credentials may reach more records than the requesting person is entitled to use.

Role-based rules assign permissions through roles. Attribute-based access control evaluates requester, resource, operation and environmental characteristics. An application can express tenant, purpose and recipient constraints through such attributes; their values must be trustworthy and current.

Access and disclosure need separate gates

Example

Permitted retrieval can still produce a prohibited disclosure.

The application checks current authority before retrieval and before release. Either gate can stop the operation; the model supplies neither permission.
Read the diagram as text
  • Requester and purpose.
  • Access gate.
  • Retrieve and draft.
  • Recipient gate.
  • Authorized recipient.
  • Stop without disclosure.
  • Requester and purposeAccess gate: Control: current authority.
  • Access gateRetrieve and draft: Control: access permitted.
  • Access gateStop without disclosure: Control: denied or unknown.
  • Retrieve and draftRecipient gate: Data: proposed response.
  • Recipient gateAuthorized recipient: Data: disclosure permitted.
  • Recipient gateStop without disclosure: Control: denied or unknown.
  • Narrow retrievalSeparate indexes can restrict candidate documents by audience or version. Selection is useful, but choosing an index does not establish complete authorization. Search and Retrieval covers candidate-selection mechanics.
  • Preserve delegation contextA multi-service action must carry the original user’s authorization context. A verifiable delegation token can support attribution, but its existence alone does not establish revocation handling or appropriately narrowed privileges.

Least privilege grants only necessary authority. Complete mediation checks every protected access, including cache reads, exports, tool calls and recovery paths. Cached decisions must account for revocation. When required authority cannot be established, deny the operation. AI Security explains enforcement against attacks; model willingness is never the authorization boundary.

Approval and authoritative record changes

A system of record is the designated authoritative source for a business fact. A data contract records the producer–consumer agreement around that information. Integration into existing work explains why field meanings, identifiers, ownership and support responsibilities matter.

Permission to summarize a customer conversation does not include permission to update contact details. Read, edit, delete and copy are distinct operations.

Approval survives as history, not permission

Example

Changed record state defeats the reviewed precondition.

1 / 3 · Proposal

C1 targets v1.

Change C1 never commits. Approval against v1 remains historical evidence after the current record becomes v2; a revised proposal needs renewed checks.
Read the diagram as text
  • Contact change C1.
  • Proposed against v1.
  • Approved against v1. Historical approval.
  • Current record v2.
  • Commit rejected.
  • Contact change C1Proposed against v1: Proposal.
  • Proposed against v1Approved against v1: Reviewer authorizes.
  • Approved against v1Commit rejected: Precondition no longer holds.
  • Current record v2Commit rejected: Version mismatch.
  1. Proposal. C1 targets v1. Active: Contact change C1, Proposed against v1. New: Contact change C1, Proposed against v1.
  2. Approval. Review binds to v1. Active: Contact change C1, Proposed against v1, Approved against v1. New: Approved against v1.
  3. Rejection. v2 prevents commitment. Active: Contact change C1, Proposed against v1, Approved against v1, Current record v2, Commit rejected. New: Current record v2, Commit rejected.
  • Bind consequential approvalWhere independent approval is required, enforce both authorities at the protected write. Bind approval to the specific proposal; separate credentials alone do not establish separation of duties.
  • Do not expand scope silentlyIf execution discovers a missing permission, request additional authorization or stop that operation. Dynamic scope negotiation is a mechanism for obtaining authority, not permission to invent it.

An HTTP update carrying If-Match: "v1" must not execute when the current strong entity tag differs. This prevents stale writes, not unauthorized or factually incorrect ones.

Keep an AI-inferred attribute distinguishable from a confirmed fact. Correction requires an accountable disposition and a new authorized update; a historical error can remain accurately recorded alongside its correction.

Lineage and proportionate audit evidence

Provenance describes origin and history; lineage records derivation relationships. W3C PROV distinguishes information entities, transforming activities and responsible actors. Participation in one activity does not itself establish derivation.

Dataset versions locate a collection; record-level relationships identify particular contributors. Retrieval-use graphs can reveal which chunks appeared in conversations, but frequent use establishes neither correctness nor permission. Citation links also require preserving the backend path to source facts.

Derivation differs from responsibility

Example

Correction follows contributing artifacts, not shared ownership alone.

Source S1 contributes through summarization and retrieval to reply R1. The owner is responsible for summarization, not a data input. Recorded edges remain assertions requiring verification.
Read the diagram as text
  • Source S1 v1.
  • Summarize.
  • Summary M1.
  • Build retrieval representation.
  • Index entry I1.
  • Draft reply.
  • Reply R1.
  • Responsible owner.
  • Source S1 v1Summarize: Used.
  • SummarizeSummary M1: Generated.
  • Source S1 v1Build retrieval representation: Used.
  • Build retrieval representationIndex entry I1: Generated.
  • Summary M1Draft reply: Context.
  • Index entry I1Draft reply: Retrieval contribution.
  • Draft replyReply R1: Generated.
  • Responsible ownerSummarize: Responsible for.
A proposed action record separates attribution from payload storage.
EvidenceInvestigation purpose
Source and output identifiers; versions; transformationLocate the information used and produced.
Purpose, policy version, requester, approver, recipient and outcomeReconstruct the authorization and resulting action.
Restricted payload referencePermit investigation without duplicating content into ordinary developer logs.

Separating orchestration events from versioned sensitive objects lets developers inspect execution without receiving payload access. References and metadata still need privacy review. An immutable event log preserves recorded history; it cannot supply events that were never captured.

Restrict and review audit access, protect integrity, sanitize untrusted event fields, and apply disposal rules to extracts and backups. Avoid copying credentials or sensitive payloads for convenience. Capture policy governs which diagnostic information is emitted.

  • Centralize consistent attributionUber’s described Model Gateway combines request attribution, policy middleware and audit-session capture. Such a shared boundary can standardize records; the example does not specify stored-trace retention or complete downstream coverage.

Vendor and model-processing obligations

A data processing agreement documents processing obligations. Under EU controller–processor guidance, a subprocessor is another processor engaged in the chain. Authorization, change notification, audit information and return-or-delete arrangements require explicit handling. Contract terms do not establish fulfillment.

No-training commitments and no-retention commitments answer different questions. Anthropic’s feature-specific retention documentation, checked August 29, 2026, distinguishes storage-dependent features from zero-retention eligibility. It separately describes training permission and hosting relationships. Those distinctions must be checked for the actual product and configuration.

Review dimensionRequired distinction
Training and feature storageDocument permitted reuse separately from storage required by enabled features.
Monitoring and support accessIdentify processing purpose, authorized people and retained information.
Configuration and evidenceCheck the selected feature and route; general provider statements do not prove account settings.
ExitWithin the EU relationship, return or delete as directed; address copies and applicable storage exceptions.

Data residency concerns where information resides; cross-border transfer concerns movement or access across national boundaries. A VPC—an isolated cloud network—does not establish government-cloud or on-premises availability. The managed-RAG deployment discussion illustrates that distinction, not current vendor suitability.

Self-hosting transfers more software, patching and availability work to the operator. Managed services transfer some infrastructure work to providers. Neither arrangement transfers the application’s responsibility for data classification, permissions and appropriate use; location alone cannot settle those obligations.

Retention schedules and restricted preservation

Retention is how long information remains stored or available. A cutoff starts a retention period; disposition is eventual destruction or transfer. NARA’s scheduling guidance provides a useful event-based model, not disposal authority for private applications.

Each schedule also needs an approved duration, owner, disposal action and completion evidence; these are example triggers, not prescribed periods.
Artifact classPossible triggering event
Source cases and uploaded documentsCase closure
Prompts, outputs and feedbackApproved operational purpose ends
Indexes and cachesSource eligibility ends; derived cleanup verified
Telemetry and audit extractsInvestigation or operational retention trigger
Training datasets and modelsApproved learning or deployment use ends
BackupsReplacement cycle, subject to applicable restrictions

A legal hold is a preservation obligation requiring qualified interpretation. Document its basis, covered records, owner, access restrictions and release conditions. An exception permitting preservation does not automatically permit continued ordinary processing.

For valid erasure requests, UK guidance addresses backups as well as live systems. Copies awaiting scheduled overwrite must remain beyond use, not repurposed. Explain that status to the person; ending active access is different from completed destruction.

Lifecycle fulfillment across derivatives

Stopping a purpose, revoking authority, correcting a fact and deleting a copy are distinct operations. Keep a destination ledger so successful work at one store cannot conceal unfinished work elsewhere.

For an accepted request, record each destination’s action, owner, evidence and status: pending, verified, excepted or unresolved.
DestinationCompletion evidence
Source and derived recordsAccepted correction and its downstream disposition
Indexes, caches and exportsRemoved content is unavailable through each affected serving path
Vendor copiesDestination-specific fulfillment evidence, not just a submitted request
Evaluation and training examplesAffected versions identified and reuse stopped or changed
Backups and restorationApplicable restriction remains effective when restoration is tested

Keep the marker until removal is verified

Example

Source deletion and search invisibility are separate facts.

1 / 4 · Available

Both copies remain available.

For a supported one-to-one indexing arrangement, S1 remains observable until I1 removal is verified. Superseded availability states disappear; artifact identities persist.
Read the diagram as text
  • Source S1.
  • Index entry I1.
  • S1 live.
  • I1 searchable.
  • S1 soft-deleted; marker retained.
  • I1 removal verified.
  • S1 purged.
  • Source S1S1 live: State.
  • Source S1S1 soft-deleted; marker retained: State.
  • Source S1S1 purged: State.
  • Index entry I1I1 searchable: State.
  • Index entry I1I1 removal verified: State.
  1. Available. Both copies remain available. Active: Source S1, Index entry I1, S1 live, I1 searchable. New: Source S1, Index entry I1, S1 live, I1 searchable.
  2. Mark. I1 remains searchable. Active: Source S1, Index entry I1, S1 soft-deleted; marker retained, I1 searchable. New: S1 soft-deleted; marker retained.
  3. Process and verify. Indexer processes the marker; query checks removal. Active: Source S1, Index entry I1, S1 soft-deleted; marker retained, I1 removal verified. New: I1 removal verified.
  4. Purge. Remove S1 after downstream verification. Active: Source S1, Index entry I1, S1 purged, I1 removal verified. New: S1 purged.

Azure’s blob deletion policy needs an observable marker and sufficient retention for indexer outages. Its soft-delete policies exclude one-to-many indexing; those derived entries require explicit deletion.

Deletion acknowledgement and search visibility can differ. Elasticsearch exposes a refresh boundary; deleted-version retention is also finite. Delayed ingestion therefore needs durable deletion markers or version checks beyond that window to prevent obsolete records from returning.

  • Retry without recreating dataA deletion ledger must preserve outstanding destinations across failures. Retry contracts need stable operation identity and defined retention; merely logging an identifier does not make a receiving API idempotent.
  • Separate model influenceMachine unlearning attempts to remove selected training influence. A verification study on image classifiers showed that dishonest providers could pass the studied checks while retaining information. Stored-record deletion therefore requires different evidence from claims about learned behavior.

Deleting a stored artifact cannot retract an earlier disclosure. Report verified destinations, restricted exceptions and unresolved effects separately; the accountable owner decides the remaining response rather than declaring universal completion.

Assurance and operating changes

A privacy impact assessment examines processing, effects on people, alternatives, mitigations and residual risk—the risk remaining after controls. A DPIA is a jurisdiction-specific data protection impact assessment; the ICO guidance distinguishes legal triggers from wider good practice and calls for reassessment when processing changes.

Assign an owner to each rule and retain the observation, gap and operating decision.
BoundaryFocused verification
Retrieval and recipientAttempt unauthorized retrieval and a recipient change; inspect actual disclosure.
WithdrawalConfirm the withdrawn purpose stops across affected processing.
Record changesCheck rejected stale proposals and authorized correction outcomes.
DeletionTest cache, export and restoration paths independently.
Vendor changesReview changed processing conditions before extending operation.

Changing product context changes the decision. Workday’s described progression from help articles to page-aware assistance introduced compensation information with greater sensitivity. Reusing an assistant interface did not preserve the original data-risk boundary.

When a boundary fails, contain affected processing, preserve necessary restricted evidence, investigate scope and correct the cause. Resume only the paths whose restoration criteria have been verified with their owners. Passing fresh-request checks cannot justify an untested cache or restore path. AI Security covers the broader response discipline.

Open questions

  1. Derivative anonymity remains difficult to establish across changing access conditions. Embedding reconstruction and contextual attribute inference expose different risks; progress requires evaluations covering the actual encoder, accessible representations and plausible auxiliary information, rather than treating a transformation label as evidence.

  2. Disputed AI attributes need context-sensitive correction rules. The boundary between factual assertion and recorded opinion changes the response, while either can affect a person. Progress would include documented dispositions tied to what each record asserts and the consequences of its use.

  3. Deletion remains vulnerable to delayed replay. Finite search-engine deletion history can outlive neither every queue nor every restoration source. Progress requires demonstrated rejection of obsolete updates after ordinary deletion metadata expires, with explicit coverage of restoration paths.

  4. Verifying removal of training influence remains an assurance problem distinct from deleting examples. The studied verification methods can be fooled by dishonest providers; extending trustworthy checks to deployed language models requires evidence beyond the image-classification settings already tested.

Follow the curated reading path through the speakers and demonstrations behind this entry.

Explore more talks

The rest of the library, beyond the curated path. Cited talks support this entry; reviewed transcripts were processed in full. Metadata candidates have not been reviewed as sources or verified as topic members.

5 matching talks

TalkSpeakerEventYear
Sumit AgarwalAI Engineer World's Fair 20242024
Ofer MendelevitchAI Engineer Summit 20252025
Lovina DmelloAI Engineer World's Fair 20262026
Andreas KolleggerAI Engineer World's Fair 20252025
Nina Lopatina, Rajiv ShahAI Engineer World's Fair 20252025

References

Coverage and source review
Processed transcripts
11 processed in full · 6 in the curated path
Automated source review
Passed
Metadata candidates
0 unreviewed; not verified topic membership
Corpus version
1bd8e407b26a07b33815594e1b2db5f41827119a2b3cb6fbf240f9fc571fc767

Automated review checks source support; it is not publication approval.

A synthesis of selected conference talks and technical references. Citations link to the source material; they do not imply that every talk on this subject is included.

  1. NIST Privacy Framework 1.0: lifecycle and minimized audit evidence

    Core ID.IM-P; GV.PO-P1; CT.PO-P; CT.DM-P5/P8; CM.AW-P6; PR.AC-P; PR.DS-P3.

  2. Regulation (EU) 2016/679: definitions, processing principles and correction

    EU GDPR Articles 4–6, 16 and 19; terminology and distinct handling obligations.

  3. NIST AI Risk Management Framework 1.0

    GOVERN 2, 3.2 and 5; MAP 1 and 5; MEASURE 1.2, 2.2–2.11 and 3; MANAGE 1 and 4.

  4. LLM Safeguards: Security, Privacy, Compliance, Anti-Hallucination

    Query data sources using the requesting user's proper role; organizing documents does not solve authorization.

  5. LLM Safeguards: Security, Privacy, Compliance, Anti-Hallucination

    Retrieved internal text can expose personal information through generated responses; filter sensitive content before it reaches the LLM.

  6. NIST SP 800-162: Guide to Attribute Based Access Control

    Sections 2.2–2.3; ABAC definitions and basic decision/enforcement model.

  7. EDPB Opinion 28/2024 on personal data in AI models

    EU GDPR opinion, sections 3.2–3.4; model anonymity, lawful basis and development-to-deployment accountability.

  8. ICO: Data protection impact assessments

    UK regulatory guidance; DPIA definition, process and review checklists.

  9. U.S. Copyright Office: Copyright and Artificial Intelligence, Part 3—Generative AI Training

    US copyright analysis; May 2025 pre-publication report, acquisition discussion, sections III, IV.A.1, V.C and VI. Establishes a copyright-specific permission question independently of personal-data rules.

  10. LLM Quality Optimization Bootcamp

    Real production prompts paired with responses from a stronger LLM are presented as a middle ground between fully human and fully synthetic training data.

  11. Beyond Memorization: Violating Privacy via Inference with Large Language Models

    Version 2, introduction and sections 3.1, 4–5. Attacker access comprises user-written texts and model prompting; attributes include location, age and income.

  12. The Adversarial Path to the Personal Assistant

    The speaker recommends answering from processed profiles where possible and consulting raw records only when needed.

  13. Text Embeddings Reveal (Almost) As Much As Text

    EMNLP 2023 paper; reconstruction method, tested encoders and clinical-note experiment.

  14. LLM Quality Optimization Bootcamp

    The speaker's crawl, walk, run framework starts with prompt engineering, adds Retrieval Augmented Generation (RAG) for missing context, and considers fine-tuning for persistent failures on narrowly defined tasks.

  15. OpenTelemetry: Handling sensitive data

    Official security guidance, implementer responsibility, data minimization, and Collector processors.

  16. Persistence — LangGraph

    Official documentation; checkpointer versus store and persistence failure modes.

  17. EDPB Guidelines 07/2020 on controller and processor

    EU GDPR guidance, Part I and Part II sections 1.3.7–1.6.

  18. Open Data Contract Standard v3.0.0

    First-use data-contract explanation and an authoritative checklist for integration conversations.

  19. OWASP Access Control

    OWASP; overview, least privilege, centralized checks and protected-resource examples. AI application is an engineering inference.

  20. Directive (EU) 2016/943: lawful and unlawful handling of trade secrets

    EU Directive Articles 2–5. A trade secret must be secret, commercially valuable because of secrecy, and subject to reasonable protective steps.

  21. Regulation (EU) 2016/679: notices and conditional individual rights

    EU GDPR Articles 12–15, 18 and 21; supplements the reused definitions note without replacing it.

  22. EDPB Guidelines 05/2020 on consent

    EU GDPR consent guidance, sections 5.1–5.2; consent evidence and purpose-specific withdrawal.

  23. ICO: Right to rectification

    UK guidance on disputed facts, opinions and request handling. Applying this distinction to an inferred attribute requires assessing what the record actually asserts.

  24. ICO: Right to erasure

    UK guidance; backup handling and exceptions to erasure.

  25. NIST AI RMF 1.0: accountability, appeals and override

    GOVERN 2.1–3 and 3.2; MEASURE 3.3; MANAGE 2.4 and 4.1–3; Appendix C. Record fields are an implementation proposal.

  26. Graph Intelligence: Enhance Reasoning and Retrieval Using Graph Analytics

    Separate vector indices can constrain the candidate corpus to a use case, permitted audience, or customer software version.

  27. LLM Quality Optimization Bootcamp

    Separating PII extraction from text replacement allows the same model output to support multiple redaction policies.

  28. Re-identification risk in released recommendation histories

    Attack model and matching method; Netflix Prize experiments and section 5 on public IMDb auxiliary information.

  29. OpenTelemetry: Handling sensitive data

    Sections 'Your responsibility', 'Sensitive data considerations', 'Data minimization', and 'Protecting sensitive data', including hashing limitations.

  30. AWS: Shared Responsibility Model

    Customer responsibilities for EC2 and abstracted services; Shared Controls; Applying the AWS Shared Responsibility Model in Practice.

  31. CIAM for AI: Authn/Authz for Agents — Michael Grinich, CEO of WorkOS

    Delegation chains pass verifiable authorization step by step while carrying the original user's authorization context.

  32. The Protection of Information in Computer Systems: Basic Principles

    Section I.A.3, Design Principles, especially fail-safe defaults, complete mediation, and least privilege; section I.B, isolation mechanisms.

  33. The Protection of Information in Computer Systems: Basic Principles

    Section I.A.2 Controlled sharing and protected subsystems; I.A.3 principles c, e and f; I.B discussion of principals.

  34. CIAM for AI: Authn/Authz for Agents — Michael Grinich, CEO of WorkOS

    GNAP, Grant Negotiation and Authorization Protocol, is proposed as a framework for dynamically negotiating scopes during execution.

  35. RFC 9110: HTTP Semantics

    Section 13.1.1; conditional mutations and lost-update prevention.

  36. PROV-DM: The PROV Data Model

    W3C Recommendation, sections 2 and 5; derivation, attribution, association and delegation.

  37. Graph Intelligence: Enhance Reasoning and Retrieval Using Graph Analytics

    Represent conversations and their retrieved chunks as a graph to expose repeated source use and shifts between knowledge areas.

  38. The Hidden Costs of Building Your Own RAG Stack — Ofer Mendelevitch, Vectara

    Clickable evidence in the UI requires the backend to preserve provenance through the response flow.

  39. Why Your Enterprise Tech Stack Isn't Ready for AI Agents - And What to Build Instead

    Design the audit record to preserve actions, data access, and the authorization behind each action, rather than treating diagnostic logs as sufficient.

  40. Why Your Enterprise Tech Stack Isn't Ready for AI Agents - And What to Build Instead

    Separate orchestration events from immutable, schema-driven objects containing sensitive data.

  41. Why Your Enterprise Tech Stack Isn't Ready for AI Agents - And What to Build Instead

    Use a unified, immutable event log as the source of truth, accepting more complex reads in exchange for reconstructable history.

  42. OWASP: Logging Cheat Sheet

    Data to exclude; Event collection; Protection; Disposal of logs.

  43. Agentic SDLC at Uber - Building Blocks for Uber’s Software Factory

    Uber places request attribution, policy middleware, and audit-session capture in a shared Model Gateway.

  44. Claude API and data retention

    Provider documentation as retrieved on the verification date; retention approach and arrangement boundaries.

  45. Forget RAG Pipelines—Build Production-Ready AI Agents in 15 Minutes

    The deployment options described do not establish support for government-cloud or custom on-premises requirements.

  46. NARA: Guide to Inventorying, Scheduling and Disposition of Federal Records

    Pages 37–39; cutoff instructions, retention periods and disposition terminology.

  47. Elasticsearch: Delete a document

    Optimistic concurrency control; Versioning; Routing; refresh query parameter.

  48. LLM Quality Optimization Bootcamp

    A fine-tune is not a one-time deployment because production data can drift, requiring dataset updates and renewed training and evaluation.

  49. NIST SP 800-61r3: incident response and verified recovery

    April 2025 final revision; RS.MA, RS.AN-06/07, RS.MI, and RC.RP-01 through RC.RP-06.

  50. Changed and Deleted Blobs — Azure AI Search

    Opening deletion-policy warning; Prerequisites; Native blob soft delete requirements and retention; Reindex undeleted blobs; Soft delete strategy using custom metadata. Limitations: one-to-many indexing.

  51. Making retries safe with idempotent APIs — Amazon Builders' Library

    Primary engineering account; client request identifiers, atomicity, semantic equivalence, and late requests.

  52. Verification of Machine Unlearning is Fragile

    Version 1; proposed adversarial methods and experiments on MNIST, CIFAR-10 and SVHN.

  53. Build Dynamic Products, and Stop the AI Sideshow

    Moving from help-article generation to contextual task assistance can expose substantially more sensitive data, including employee pay and compensation information.

  54. LLM Safeguards: Security, Privacy, Compliance, Anti-Hallucination

    Surface prompt-injection flags and patterns of sensitive-data submission while minimizing prompt and completion retention.

  55. Build Dynamic Products, and Stop the AI Sideshow

    Cross-suite assistance and proactive policy responses require coordination across product teams beyond the scope of a single feature or SKU.