What happens to your data when a supplier builds your software.
By Zain M · Updated 14 September 2026 · 14 min read
When a supplier builds or runs software for you, you remain the data controller and the supplier your processor. UK GDPR Article 28 requires a written contract, and the ICO lists eight clauses covering instructions, security, sub-processors, rights requests, breach assistance, deletion and audit. Any AI model provider the software calls is a sub-processor and belongs in the same agreement.
The distinction that decides everything
You decide why and how personal data is used, which makes you the controller and puts the legal obligation on you. A supplier acting on your instructions is a processor. That remains true even when the supplier knows far more about the technology than you do, and it remains true when the supplier is a one-person studio or a large consultancy. The ICO’s detailed guidance on controllers and processors exists to settle exactly this question, and its answer turns on who decides the purpose and the means, not on who holds the servers.
It matters because the regulator will come to you, not to them. Which is precisely why the agreement between you needs to be written rather than assumed. The ICO puts it plainly: whenever a controller uses a processor to process personal data on their behalf, a written contract needs to be in place between the parties, and if the processor uses a sub-processor, it needs a written contract with them too.
Software that calls an AI model adds a layer most agreements from before 2023 never contemplated. Every request that leaves your system and reaches OpenAI, Anthropic, Google or Microsoft is processing by a sub-processor. Their terms decide whether your data trains a model, how long it is kept and in which country it is stored, and they differ between the consumer, business and API versions of the same product. This guide covers the contract, the transfer rules, the provider terms and the practical arrangements, in that order.
What Article 28 says the contract must contain
The ICO publishes a checklist for the controller-processor contract. It has two parts: the description of the processing, and the eight clauses the contract must include. A supplier’s standard data processing agreement should map onto this table line by line; if it does not, ask for the missing lines before signing. The right-hand column is the question to put to the supplier for each line, because the clause is only worth having if the supplier can say how it is met in practice.
| What the ICO checklist requires | What to ask the supplier |
|---|---|
| Subject matter, duration, nature and purpose of the processing; type of personal data; categories of data subject; the controller’s obligations and rights | A schedule naming the data, the people it relates to, and what the system does with it |
| The processor acts only on the controller’s documented instructions, unless required by law | Are the instructions written down, and is model training, analytics or "service improvement" excluded unless you say otherwise? |
| People processing the data are subject to a duty of confidence | Who on the team and among contractors can see live data, and under what agreement? |
| Appropriate measures to ensure the security of processing | Encryption, access control, logging, and how test environments avoid live personal data |
| Sub-processors only with prior authorisation and under a written contract | The named list, including every AI model provider and hosting region, and your right to object to changes |
| Measures to help the controller respond to individuals’ rights requests | How a subject access request or erasure request is fulfilled across the database, backups and the model provider |
| Assistance with security, breach notification and DPIAs | A breach notification commitment in hours, and help completing the DPIA the ICO expects for most AI |
| Delete or return all personal data at the end, and delete existing copies unless the law requires storage | What is deleted, when, from where, and what evidence you receive |
| Submit to audits and inspections | Can you, or an auditor you appoint, inspect the arrangements? |
From the ICO’s Guide to UK GDPR, contracts checklist, read 14 September 2026. The processor must also give the controller the information it needs to show both are meeting Article 28.
Where the model provider sits, and what it does with your data
The provider of the AI model is a sub-processor, and its terms are the part of the chain a buyer most often never reads. The difference that matters is between the consumer product and the business or API product with the same name: the consumer version may train on what it is sent unless the user opts out, and the business version does not by default. A build should only ever use the second kind. Read on the providers’ own pages in September 2026, the position is this.
| Provider and product | Used to train models? | Retention stated | Region options | Source, date read |
|---|---|---|---|---|
| OpenAI API, ChatGPT Business, Enterprise, Edu | No, by default; API data has not been used for training since 1 March 2023 | API inputs and outputs up to 30 days for abuse monitoring; zero data retention on approval for eligible endpoints | Data residency in the United Kingdom, Europe (EEA and Switzerland), US and others; 10% uplift for models released on or after 5 March 2026 | OpenAI enterprise privacy and developer docs, Sep 2026 |
| OpenAI ChatGPT for individuals (Free, Plus, Pro) | Yes, unless you turn off "Improve the model for everyone" or use Temporary Chat | Governed by the consumer privacy policy | None to select | OpenAI help centre, Sep 2026 |
| Anthropic API (Commercial Terms) | No: "Anthropic may not train models on Customer Content from Services"; the customer owns outputs | Per the commercial terms and retention settings | Inference geo "global" or "us"; US-only inference at 1.1x price; workspace data-at-rest geo is US only at present | Anthropic commercial terms, effective 17 Jun 2025; data residency docs, Sep 2026 |
| Anthropic Claude Team and Enterprise | No model training on your content by default | Per customer agreement | As above | claude.com pricing, Sep 2026 |
| Anthropic Claude consumer (Free, Pro, Max) | May be used unless you opt out in settings | Up to 5 years de-identified if you allow training; deleted chats removed from back-end within 30 days | None to select | Anthropic privacy policy effective 10 Sep 2026; retention article 1 Jul 2026 |
| Google Gemini API, paid tier | No: Google "doesn’t use your prompts … or responses to improve our products" | Logged for a limited period solely to detect prohibited use | Vertex AI lists London (europe-west2) among regions for Gemini and Claude models | Gemini API terms effective 23 Mar 2026; Vertex locations page, Sep 2026 |
| Google Gemini API, unpaid tier | Yes: content is used "to provide, improve, and develop Google products and services" | As stated in the terms | None to select | Gemini API terms, Sep 2026 |
| Microsoft Copilot and Copilot Chat (work accounts) | No: prompts, responses and Graph data are not used to train foundation models | Your tenant’s retention policies apply | EU Data Boundary supported; Anthropic models within Copilot are currently excluded from it | Microsoft Learn, 29 May 2026 |
Vendor statements, read on the vendors’ own pages; they change, and the date read is the date to check against. Regions are options that a builder must switch on; the default is often global.
Leaving the UK: adequacy, the IDTA and the Addendum
If personal data is sent, or made accessible, to a separate organisation outside the UK, the ICO calls it a restricted transfer, and every restricted transfer must be covered by one of three things: UK adequacy regulations for the destination country, an appropriate safeguard, or an exception. The safeguards the ICO names are its International Data Transfer Agreement (the IDTA), the International Data Transfer Addendum to the EU standard contractual clauses (the Addendum), and UK binding corporate rules. If you rely on a safeguard you must also complete a transfer risk assessment to check that the standard of protection for people’s information is not materially lower after the transfer. If none of the three routes is available, the transfer must not be made.
For a software build the practical questions are where the database sits, where backups sit, and where each model provider processes the request. A supplier who hosts in London and calls a model with a global routing default has made a restricted transfer whether they meant to or not. The provider pages above show what can be pinned: OpenAI offers UK and European residency at a 10 per cent uplift on newer models; Google’s Vertex AI lists London as a region for Gemini and for Claude models; Anthropic’s first-party API currently offers US or global inference and US-only storage; Microsoft offers the EU Data Boundary for Copilot. Each is a setting, and each costs something, so it belongs in the scope.
The Data (Use and Access) Act 2025 restated the test for third-country transfers as whether protection would be "not materially lower", which is the language now used in the ICO’s transfer risk assessment guidance. It also removed the four-year review requirement on adequacy decisions. Neither changes what a small buyer has to do: know where the data goes, and have a mechanism for each hop.
What the Data (Use and Access) Act 2025 changed
The Act received Royal Assent on 19 June 2025 and, per the ICO’s guidance page updated on 19 June 2026, all of its data protection provisions are now in force, having been phased in between June 2025 and June 2026. It amends rather than replaces the UK GDPR, the Data Protection Act 2018 and PECR, and the ICO’s summary is that most of the changes offer an opportunity to do things differently rather than a new obligation.
Four changes touch a software build directly. Automated decision-making: the Act replaced Article 22 with new Articles 22A to 22D and, in the ICO’s words, opens up the full range of lawful bases you can rely on when using personal information to make significant automated decisions, including legitimate interests, so long as safeguards such as the right to human intervention remain and special category data is not involved. Recognised legitimate interests: a new lawful basis for listed purposes such as crime prevention and safeguarding that removes the balancing test.
Subject access requests: you now only have to make reasonable and proportionate searches, and the clock can be stopped while you seek clarification. Cookies: some storage and access technologies, such as those used for statistics or website functionality, can be set without consent. None of these requires a change to an existing system; each is an option a new build can be designed to use.
There is one new duty: organisations must have a data protection complaints procedure, and the ICO has published guidance on it. For a system that scores, ranks or triages people, the ADM change is the one to design around: the wider lawful basis comes with the same duty to tell people, let them contest the decision and get a human to look at it.
Account ownership, escrow and the exit
The paperwork matters less than three practical arrangements. First, the systems should sit on your own provider accounts, with the supplier given access rather than ownership: the cloud project, the domain, the model provider organisation, the messaging account. That single arrangement removes most of the ambiguity about what happens if the relationship ends badly, and it means the sub-processor relationship is between you and the provider, on terms you accepted, with the invoices in your name.
Second, the code should be in a repository you own from the first commit, not handed over at the end. Source-code escrow, where a third party holds the code and releases it on the supplier’s insolvency, is the traditional answer for large packaged software; for a bespoke build on your own accounts it is usually unnecessary, because there is nothing to release. If a supplier insists on holding the code and the accounts, escrow is the minimum, and a retainer that exists mainly because they hold the login is not support.
Third, the exit should be rehearsed on paper before it is needed: what is deleted, from where, including backups and the model provider’s retention window, and what written confirmation you receive. The ICO’s end-of-contract clause requires deletion or return at your choice; the practical test is whether the supplier can tell you, today, exactly where every copy of your data is.
A worked example: the clock, in hours
Two deadlines fall on you as controller, and both are shortened by whatever the supplier keeps for itself. A personal data breach must be reported to the ICO within 72 hours of you becoming aware of it where it is likely to result in a risk to people. If the supplier’s agreement says "without undue delay" and nothing else, and they take two days to tell you, you have 24 hours left to assess, document and notify. A supplier that commits to notifying within 24 hours leaves you 48. Ask for the number.
A subject access request must be answered within one calendar month, extendable in complex cases and now pausable while you seek clarification under the 2025 Act. If the system holds the person’s data in a database, a search index, a model provider’s 30-day abuse-monitoring log and a nightly backup, the supplier needs to be able to search all four. A supplier who can produce a person’s complete record in a day makes the month easy; one who has to be asked what the system stores does not.
Costs to put against those: a written DPA is a short document for a small engagement, and pinning a model provider to a UK or European region adds about 10 per cent to model spend at OpenAI’s published uplift and 10 per cent at Anthropic’s for US-only inference. On a system whose model bill is tens of pounds a month, the residency premium is pounds. The expensive version of this guide is the one where none of it was done.
The checklist
Score a proposed supplier against this before signing. Every "no" is a conversation, not necessarily a rejection; a supplier who answers all of them from memory is one who has done this before. The evidence column matters more than the answer, because "yes" is free and a screenshot of the provider settings, a login to the repository, or a dry run of a rights request on a test record is not.
One item on the list is easy to check without asking: the ICO keeps a public register of data controllers, searchable by name or number, and a supplier that processes personal data for a living and is not on it has already told you something about how it treats the rest of the list. The registration number should appear on the supplier’s privacy notice; ours is at the foot of every page.
| Item | What good looks like | Evidence to ask for |
|---|---|---|
| Written DPA offered before commitment | Their standard agreement, mapped to the ICO checklist | The document itself |
| Roles stated | You controller, they processor, providers sub-processors | The DPA’s parties clause |
| Sub-processor list | Every provider named, with region and training position | A schedule you can object to changes in |
| Model training excluded | Business or API terms only; no consumer accounts in production | Screenshot of the organisation settings, or the terms they signed |
| Data location | Database, backups and model inference regions named | The cloud region and the provider residency setting |
| Transfer mechanism | Adequacy, IDTA or Addendum for each hop outside the UK | The mechanism named per provider |
| Breach notification | A commitment in hours | The clause |
| Rights requests | A described route to search and delete across every store | A dry run on a test record |
| DPIA | Drafted with you before go-live where personal data is processed by AI | The document |
| Account ownership | Cloud, domain, model and messaging accounts in your name | Your login |
| Code ownership | Repository in your organisation from the first commit | Your login |
| Exit | Deletion or return, from every store, with written confirmation | The clause and the runbook |
| ICO registration | Supplier registered as a data controller for its own business | Registration number, checkable on the ICO register |
What Augustova’s data processing agreement contains
For transparency, since we are asking you to demand this of every supplier: Augustova Limited is registered with the ICO as a data controller under number ZC152144, and our standard DPA follows the ICO checklist above clause by clause. It names us as processor and you as controller, lists every sub-processor by name and region, excludes model training by using business and API terms only, commits to breach notification within 24 hours of our becoming aware, and provides for deletion or return of all personal data at the end of the engagement with written confirmation.
The practical arrangements are the ones described in this guide because they are ours: systems on your own provider accounts with access granted to us, code in your repository from the first commit, model providers pinned to a UK or European region where the provider offers one, and a DPIA drafted with you before anything touches live personal data. We will send the DPA before you commit to anything, so you can read the terms rather than take them on trust.
The residency and no-training settings are recorded in the handover document with a screenshot of each provider’s organisation settings on the day of go-live, because a setting that was correct at launch and silently changed by a provider’s new default is the failure mode we have seen most often in systems we have been asked to review. The handover also lists every place a person’s data can be found, so that a rights request can be answered from the document rather than from memory.
Method and sources
Every regulatory statement above was read on the ICO’s or gov.uk’s pages on 14 September 2026: the ICO’s contracts checklist and controllers-and-processors guidance, its international transfers guide, its Data (Use and Access) Act page as updated on 19 June 2026, and DSIT’s factsheet on the Act’s UK GDPR changes. Provider positions on training, retention and regions were read on the providers’ own legal and documentation pages on the same day, with effective dates given where the page states one; these are vendor statements and change without notice, so check the date. The 72-hour and one-month deadlines are UK GDPR Articles 33 and 12. Nothing here is legal advice; it is what the sources say, arranged for a buyer.
Common questions
Do we need a data processing agreement for a small project?
If any personal data is involved, yes. There is no size threshold: the ICO says whenever a controller uses a processor there must be a written contract, and it lists eight clauses the contract must contain. It is also a short document for a small engagement, a few pages for a single-process build, so the cost of doing it properly is low and the cost of not having one is carried entirely by you.
Is the AI model provider a sub-processor?
Yes, if the software sends personal data to it. That means the provider must be named in your agreement, you must be able to object to a change of provider, and the provider’s own terms on training, retention and region become part of your compliance. Business and API terms at OpenAI, Anthropic and Google exclude training by default; consumer terms do not.
Does OpenAI train on data sent through the API?
No, by default. OpenAI states that data sent to the API has not been used to train its models since 1 March 2023 unless you opt in, and the same applies to ChatGPT Business and Enterprise. Consumer ChatGPT is different: content may be used unless you turn off "Improve the model for everyone" or use Temporary Chat.
Can our data be kept in the UK?
Often, but it is a setting, not a default. OpenAI offers United Kingdom data residency for eligible models at a 10 per cent uplift on models released on or after 5 March 2026. Google’s Vertex AI lists London (europe-west2) as a region for Gemini and Claude models. Anthropic’s own API currently offers US or global inference only.
What is the IDTA?
The ICO’s International Data Transfer Agreement, one of the appropriate safeguards under Article 46 of UK GDPR for sending personal data to a country without UK adequacy regulations. The alternative is the Addendum to the EU standard contractual clauses, and large groups can use binding corporate rules. Whichever safeguard is used, the ICO says a transfer risk assessment must be completed as well.
Who is liable if the supplier causes a breach?
As controller you carry the primary obligation to the regulator and to the individuals affected, and the 72-hour notification clock runs from when you become aware. Processors have direct obligations of their own under UK GDPR, but the ICO comes to you first. That is why the breach notification commitment in the supplier’s agreement should be a number of hours, not "without undue delay".
Should the supplier host the system on their own accounts?
Prefer your own accounts with access granted to them. It removes ambiguity about ownership, makes an exit straightforward, puts the sub-processor relationship and the invoices in your name, and means the data never depends on a relationship continuing. If a supplier insists on holding the accounts and the code, source-code escrow is the minimum, and you should ask why.
What did the Data (Use and Access) Act 2025 change for a software build?
Per the ICO, all its data protection provisions are in force as of June 2026. The main effects on a build are a wider set of lawful bases for significant automated decisions, with safeguards retained; a new recognised legitimate interests basis; reasonable and proportionate searches for subject access requests; and a requirement to have a data protection complaints procedure.
Do we need a DPIA for AI software?
The ICO says that in the vast majority of cases the use of AI will involve processing likely to result in a high risk to individuals, which triggers the legal requirement for a DPIA. Where you assess a use as not high risk, you still need to document how you reached that view. Ask the supplier to draft it with you.
Ask for our data processing agreement
We will send our standard one before you commit to anything, so you can see the terms, the sub-processor list and the regions rather than take them on trust.