What data do AI labs buy from companies?

· By Orca

This guide is for owners and managers of UK companies who have heard that AI labs pay for company data and want to know what that means in practice. The short answer: labs want records of how work really gets done, not your customers’ details.

What do AI labs actually want from a company?

AI models have read much of the public internet. They have seen far less of how real companies do work, step by step, with the outcome attached: the procedure, the exception, the escalation, the check. That gap is what data programmes try to fill, and it is what operational records show.

Think of it as the difference between a textbook and a working office. A model can recite what a change-control process is. It has rarely seen a real one followed, bent, escalated and signed off. Records that show the whole sequence, including what went wrong and who decided, are scarce, and that scarcity is why programmes pay attention to them.

Prices vary by programme and by what is licensed, and nobody can promise you a figure. Our homepage estimate is illustrative only: an estimate, not an offer. For what drives value, see what company data is worth to AI.

Which of your records count, and what should you strip out?

Eight kinds of record come up in our check: procedures, knowledge base, workflow history, completed work, training material, templates, quality and review records, and decision trails. The table shows what each tells a lab, what makes it useful, and what to remove first.

Record typeWhat it shows a labWhat makes it usefulStrip out first
Procedures (SOPs, checklists, runbooks, playbooks)The intended way the work is doneCurrent, specific, followed in practice, covering many tasksClient names, internal credentials, system addresses
Knowledge base (wiki pages, how-to articles)What staff look up and how problems are explainedWritten by your team, kept up to date, consistent in stylePasted vendor manuals, client-specific pages, passwords
Workflow history (tickets, cases, task and approval trails)Work as it really happened, in order, with outcomesLong run of resolved items, clear status changes, outcome recordedCustomer names, contact details, attachments, free-text personal data
Completed work (projects, deliverables, reports)What good finished output looks likeYour own work product, with the brief it answeredAnything the client owns or marked confidential
Training material (induction guides, walkthroughs)How new people are taught the jobStep-by-step, written for real tasksStaff names, third-party course content
Templates and standard documentsYour standard structure and wordingYour own, in use, and paired with filled-in examplesBorrowed or bought templates you do not own
Quality and review records (QA checks, post-mortems)What counts as good, and how mistakes are caughtConsistent criteria applied over time, with findingsIndividual staff performance detail, named clients
Decision trails (escalation notes, sign-offs)Judgement: why one option was chosen over anotherThe reasoning is written down, not only the resultPersonal or commercially sensitive detail of named parties

The strongest sets are usually several types together: a procedure, the tickets that followed it, and the review that checked them. If you are unsure whether you may share something, read client consent and data licensing first. Two sector examples are SOPs and playbooks and IT support ticket history.

What do labs generally not want?

Labs and programmes generally steer away from raw personal data, client-confidential material, generic templates copied from the internet, and very small or one-off samples. Each carries legal risk, adds nothing the model has not already seen, or is too thin to show a repeatable pattern.

  • Raw personal data. Customer, candidate and staff details bring UK GDPR duties. Remove them before anything leaves your hands. See our privacy notice for how we treat what you tell us.
  • Client-confidential material. If a client owns it or your contract restricts it, it is not yours to license.
  • Generic templates from the internet. The model has already read them. Your own versions, used in real work, are different.
  • Tiny or one-off samples. A few files cannot show how work is repeated.

This is general information, not legal advice. Ask a solicitor about your own contracts.

Do labs buy directly from mid-sized companies?

Usually not. As far as we understand it, and as press reports describe, labs tend to reach specialist knowledge through data programmes and intermediaries rather than approaching each mid-sized company. That is why a company normally deals with a programme, not the lab itself.

For example, TechCrunch reported on 29 October 2025 that labs use an intermediary marketplace to reach industry expertise they cannot easily get directly. That article is about hiring experts, not licensing company records, so treat it as context rather than proof of how every deal works. We explain the steps in how AI data programmes work.

Has a UK company actually been paid for this?

Not that we have verified. We have not seen a deal by a UK operating company that we can confirm, so we don’t claim one. The public examples we have seen concern wound-down or US firms, and any outcome for a particular company is uncertain until a programme decides.

That is why we describe this carefully. Whether your records are of interest depends on their history, consistency and the rights you hold. The honest way to find out is a short look at what you have.

What is the next step?

Take the free three-minute check. It asks which of these record types you hold, how long and how consistently, and who owns them. We review each one by hand and reply, by WhatsApp or email, with a recommendation or an honest no.

The check is free. If you go ahead and a deal closes, we charge a success fee, agreed with you in writing before anything starts. Programmes may also pay us a referral fee, which we tell you about first; neither decides what we recommend. Details are on our disclosures page, and all guides are on the guides index.

Common questions

Do AI labs want my customer data?

Generally no. Customer and client records carry personal data and confidentiality duties. What programmes look for is how your team does the work: procedures, decisions and outcomes, with names and client detail removed first.

Is a small set of documents worth anything?

Volume and consistency matter. A handful of one-off files shows little about how work is done repeatedly. Years of consistent records from a team that follows a set process give a programme far more to work with.

Who removes the personal information?

You should do a first pass, because you know your records. We only recommend programmes that keep personal and client information out of what they take, and a programme will usually run its own checks too. We help you work out what to strip.

Does the check tell me what my records are worth?

No. The check is free and tells us what you hold; we reply with a recommendation, not a valuation. The figure on our homepage is an estimate, not an offer. Any real terms come from a programme, and we only charge a success fee if a deal closes, agreed in writing first. Programmes may also pay us a referral fee, which we tell you about first.