AI projects at insurers tend to slow down at the same point. The technical team wants to try reading claim notifications with a language model; legal and information security, quite rightly, ask "where does this data go?" That question isn't asked to stop the project but to design it properly, and most of the answer lies in the architecture.

This article isn't legal advice; work with your legal and data protection officers on any real project. The examples refer to Turkey's personal data protection law (KVKK), which is close in spirit to the GDPR. What follows is a checklist of the technical measures you can take before data reaches the model and after it does.

1. Draw the data flow first

"We use AI" isn't enough. Show on a single page which data goes to which component, from the moment a claim enters the system until it's archived:

  • Where does the data come from? (messaging apps, email, web forms, call centre notes)
  • Where is it stored, and who can access it?
  • What exactly is the text sent to the language model? Do images and PDFs go too?
  • Does the model provider store the data or use it for training? What does the contract say?

This diagram clarifies the design and gives you the basis for updating your privacy notice and data inventory.

2. Mask personal data before it reaches the model

To understand a claim, the model doesn't need the policyholder's phone number, national ID or IBAN. These fields can be detected in a rule-based layer and replaced with placeholders before the text is sent:

  • national ID → [ID], phone → [PHONE], IBAN → [IBAN], email → [EMAIL]
  • names and addresses too, where needed.

Masking should be deterministic: regular expressions and validation rules (such as the check digits of a Turkish ID number), not an instruction asking the model to "mask this". The real values stay in your system, and tasks such as policy matching are done on your side with those values.

Some fields need separate thought. A licence plate matters for a claim and can be linked to a person indirectly. Does the model really need to see the plate, or should matching happen on your side? Decide field by field.

3. Images and documents: convert to text first

Photos of registration certificates, licences and accident reports contain faces, signatures and ID numbers. Instead of sending them as images to an AI service, converting them to text on your own server (OCR) and running them through the same masking layer shrinks the exposed data considerably. The model only sees masked text.

4. Special-category data: an injury is health data

A point often missed in claims: "the other driver was injured and taken to hospital" contains health data. KVKK, like the GDPR, gives health data extra protection as a special category. In practice:

  • the purpose of processing injury information (for example, prioritising the file) should be clearly defined,
  • medical details (diagnosis, treatment) shouldn't go to the model unless necessary,
  • access to these fields should be limited to a narrower group of roles.

5. Cross-border transfer: start switched off

Many popular language model services run outside the country where the data was collected. KVKK's rules on transfers abroad changed in 2024, redefining adequacy decisions, appropriate safeguards such as standard contracts, and exceptional cases. Which route fits you is a legal assessment.

What you can do technically is turn that decision into a switch in the system: use of models hosted abroad starts disabled per insurer or agency and is enabled once the legal review is complete. While it's off, the system should keep working with a locally hosted model or with rule-based steps only.

6. Storage: encrypted, separated and time-limited

  • Encryption: claim text and personal fields should be encrypted at field level in the database.
  • Separation: in a system serving several insurers or agencies, each tenant's data should be kept apart, and that separation verified by automated tests.
  • Retention: decide how long raw model-call logs are kept; debugging logs shouldn't contain personal data at all.

7. Audit trail and explainability

Being able to show later what the AI suggested on a file and what the handler did with it matters for internal audit and for any request from the data subject. Store every suggestion with the model used, the prompt version and the masked input. Write status changes to a record that can't be altered afterwards.

8. The limit on automated decisions

KVKK gives individuals the right to object when a result against them arises solely from automated analysis, similar to the GDPR's provisions on automated decision-making. That's a strong reason for decisions with consequences, such as rejection or payment, to be made by a person with a written reason. AI prepares and suggests; an authorised person decides.

Quick checklist

  1. Is the data flow diagram ready?
  2. Are ID, phone, IBAN and email masked before reaching the model?
  3. Do images go to OCR and masking first, not to the model?
  4. Is health data handled separately?
  5. Can the use of models hosted abroad be switched off per insurer?
  6. Are personal fields encrypted and separated per tenant?
  7. Is every suggestion recorded with its model and prompt version?
  8. Do consequential decisions stay with people?

Conclusion

Data protection law isn't an obstacle to AI in insurance; it's the frame of a good architecture. When masking, local OCR, cross-border transfer off by default and traceable suggestions are designed in from the start, the conversation with legal and security moves from "why we can't" to "under which conditions we can".