I spent an afternoon recently going through invoices from three different annotation vendors a client had quietly stacked up over eighteen months, and the pattern was the same one showing up across the wider industry right now: nobody had asked who actually reviewed the labelled data before it went into a production model. That gap is no longer a quiet operational detail.
It is the subject of active regulatory attention, a wave of acquisitions, and genuine cost pressure, and the data labeling news of the past few months tells UK business owners exactly where the next compliance and budget headaches are coming from.
Why data labeling news matters to UK businesses that don’t build AI models
Most UK firms reading data labeling news assume it’s a story about Silicon Valley AI labs and has nothing to do with a 40-person accountancy practice in Leeds or a mid-sized retailer in Manchester. That assumption is wrong, and increasingly expensive. If your business uses an AI recruitment screening tool, a fraud detection system, a customer service chatbot, or even a “smart” stock forecasting plug-in, someone, somewhere, labelled the data that taught that system how to behave. AI systems generally rely on vast quantities of data, whether structured, such as financial records stored in a fixed format, or unstructured, including images, videos and text files, and the quality of how that data was tagged and reviewed feeds directly into whether the tool you’re paying for actually works, and whether you can defend its decisions if challenged.
The other reason this matters now: UK regulators have started treating the provenance of training data as a compliance question, not just a technical one. That changes how procurement teams should be reading data labeling news.
The ICO is about to write the rulebook, and data labeling sits inside it
This is the single biggest UK-specific development buried in recent data labeling news, and most business owners haven’t clocked it yet. Under the Data Protection Act 2018 (Code of Practice on Artificial Intelligence and Automated Decision-Making) Regulations 2026, which came into force on 12 May 2026, the Information Commissioner’s Office is now legally required to produce a statutory code of practice covering AI and automated decision-making. This isn’t advisory guidance you can skim and ignore. Once finalised, the code carries statutory weight under the Data Protection Act, meaning courts and the ICO must take it into account in enforcement action.
Why does that touch data labeling specifically? Because the code is expected to address transparency, bias mitigation, and documentation of how AI systems were built, which loops straight back to how training data was sourced, labelled, and quality-checked. The ICO’s draft ADM guidance, out for consultation until 29 May 2026, already makes clear that organisations need “adequate mechanisms in place to diagnose quality issues” in automated systems, and that means knowing what’s underneath the model, not just trusting the vendor’s marketing page.
I’ve sat in enough vendor calls to know most UK SMEs buying AI recruitment or credit-scoring tools have never once asked the supplier how their training data was labelled, audited, or by whom. That has to change. If your procurement process for any AI tool doesn’t include a question about data labeling and human review standards, you’re building a compliance gap into your supply chain. It’s the same verification instinct UK buyers are increasingly applying elsewhere, checking the provenance of a claim before paying for it, whether that’s a diamond certificate or a vendor’s data-quality promise.
A major acquisition shows where the real money in data labeling is going
If you want evidence the data labeling industry has matured past “cheap crowdsourced tagging,” look at January’s Handshake acquisition of Cleanlab. Handshake, a platform that started out hiring college graduates and pivoted into human data-labeling for foundational AI model companies, bought Cleanlab specifically for talent, an acqui-hire bringing in nine staff including the MIT-trained founders behind Cleanlab’s automated data-quality auditing software. Handshake has provided data for eight top AI labs, including OpenAI, and was forecasted to end 2025 at $300 million in annualised revenue run rate, with that figure reportedly climbing into the high hundreds of millions through 2026.
The detail that should matter to UK readers isn’t the valuation, it’s the stated reason for the deal. Cleanlab’s CEO told reporters that competing labeling firms were already using Handshake’s platform to source the doctors, lawyers, and scientists doing specialist annotation work, so buying the auditing layer made strategic sense rather than selling to a rival labeling shop. That’s the real shift in this market: providers are no longer competing on raw labeling volume, they’re competing on who can prove the labels are correct.
The market is growing faster than most UK procurement budgets are planning for
Market sizing reports vary wildly depending on methodology, but the direction is consistent. One widely cited estimate puts the data labeling market at USD 2.61 billion in 2026, climbing to USD 7.02 billion by 2031, a 21.94% compound annual growth rate. Manual labeling still accounted for 42.31% of the market in 2025, while self-supervised and programmatic techniques are growing faster, at a 22.16% CAGR, which tells you the human-in-the-loop model isn’t disappearing, it’s becoming more specialised and more expensive per hour for the skilled portions of the work.
That cost pressure is already visible in wage data. Specialist annotators such as lawyers, doctors, and linguists are now commanding wages reaching USD 60 per hour, which is splitting the supply side into a high-skill tier and a commodity tier. For UK businesses commissioning bespoke model fine-tuning, whether through an agency or in-house, that means budgeting for labeling work the way you’d budget for any specialist contractor, not as a rounding error on the IT line.
If you’re running a small business through this kind of technology investment cycle for the first time, the labeling and review cost should be treated as a recurring operational line, not a one-off setup fee, because models need re-labelled data every time their use case shifts.
What “human in the loop” actually means now, and why vague promises aren’t enough
A recurring theme across this year’s data labeling news is that “human in the loop” has become a phrase vendors use loosely, and regulators are starting to push back on vague usage. A 2025 industry survey found that 80% of companies emphasised the importance of human-in-the-loop machine learning for successful AI projects, but the ICO’s own recruitment research found the opposite problem in practice: many employers believed they weren’t using automated decision-making when, on closer inspection, they actually were, with human review reduced to what regulators now call a “token gesture.”
This matters directly for any UK business using AI in hiring, lending, or customer triage. The ICO’s draft guidance is explicit that a decision only avoids stricter ADM rules if a human’s involvement is active and meaningful, not a rubber stamp at the end of an automated pipeline. If your HR team is approving AI-shortlisted candidates without genuinely reviewing the underlying scoring logic or labelled training examples behind it, you may be operating under tighter legal obligations than you realise.
I’d recommend any UK firm using third-party AI hiring or scoring tools ask the vendor directly: who labelled the training data, what quality checks ran on that labelling, and what does the audit trail look like if a candidate or customer disputes a decision. If the vendor can’t answer clearly, that’s the real risk, more than any headline AI feature. UK consumers are already being told to ask exactly this kind of question before trusting an unverified supplier, the same due-diligence habit we’d recommend before buying anything online without proper checks, and AI vendors deserve no less scrutiny than any other supplier making unverified claims.
EU documentation rules are reaching UK suppliers even without a UK AI Act
The UK still has no standalone AI Act, and as of mid-2026 none is before Parliament. But that doesn’t mean UK firms are insulated. The EU AI Act’s provisions on documentation and provenance metadata for training data apply to any UK business placing AI products on the EU market, providing services to EU customers, or whose AI output affects EU residents. The EU AI Act demands provenance metadata and dataset documentation, raising compliance overhead for any organisation in scope, and most of that documentation traces back to how the underlying data, including labelled training data, was sourced and verified.
For a UK exporter or SaaS provider with EU clients, this is where data labeling news stops being abstract industry commentary and starts being a contract negotiation point. If a German or French client asks for evidence of your AI tool’s data provenance, and your answer is “the vendor handles that,” you need a better answer than that before the contract renews.
Frequently Asked Questions
What is data labeling and why does it matter for AI? Data labeling is the process of tagging or annotating raw data, images, text, audio, so an AI model can learn from it. It matters because the accuracy and fairness of any AI system depends heavily on how well, and how carefully, that underlying data was labelled.
Is data labeling regulated in the UK? There’s no standalone UK law naming “data labeling” specifically, but the ICO’s forthcoming statutory code of practice on AI and automated decision-making, required under regulations that took effect in May 2026, will indirectly govern data quality and transparency standards that touch labeling practices.
Who does data labeling for AI companies? A mix of specialist firms, including Scale AI, Handshake, TELUS Digital, and iMerit, alongside in-house teams at larger AI labs. The work ranges from low-skill image tagging to highly specialised review by doctors, lawyers, and linguists.
Why are companies acquiring data labeling and auditing startups? Acquisitions like Handshake’s purchase of Cleanlab reflect a shift in the industry from competing on raw labeling volume to competing on provable data quality and audit capability, which AI labs increasingly demand before trusting a dataset.
Does the EU AI Act affect UK companies that don’t operate in the EU? Yes, if their AI system’s output is used in the EU, if it’s placed on the EU market, or if it affects EU residents, UK businesses can fall within scope of the EU AI Act’s documentation and provenance requirements even without an EU office.
Final Thoughts
What strikes me most, watching this space closely, is how quickly “just outsource the labeling” stopped being a safe default. UK firms buying or building AI tools now need to ask harder questions about who labelled the data behind them and how that work was checked, because regulators are asking those same questions.
I’d treat every new data labeling news story this year as a prompt to revisit your own AI vendor contracts, not just an industry curiosity, and it’s worth browsing our wider blog coverage of UK buying decisions if this is the first time you’ve had to vet a supplier’s claims this closely. For the clearest sense of where UK expectations are heading, the ICO’s AI and biometrics strategy is the most reliable primary source to watch.

AlphaMarket.co.uk is a business-focused platform dedicated to helping individuals and entrepreneurs grow in the modern digital world. We provide practical insights, guides, and strategies on Online Business, Digital Marketing, Small Business, Finance, and E-commerce. Our goal is to simplify complex business concepts and deliver actionable content that helps readers start, manage, and scale their ventures effectively