Bring structure, visibility, and compliance to your transport operations

Book a demo

Summary

A practical guide to AI in logistics for mid-market European shippers. It covers how much AI European transport actually uses, which capabilities have published evidence behind them (document extraction, ETA prediction, demand forecasting) and which do not, what accuracy to demand from vendors in writing, why most AI projects stall before production, and what AI costs at mid-market scale. It also sets out what the EU AI Act and GDPR require of a shipper today and from December 2027, what not to automate, and a five-step sequence for getting started.

Most writing about AI in logistics describes an industry that does not exist yet at your scale. It names twenty use cases, quantifies none of them, and assumes a budget and a data team you do not have. Read enough of it and you come away knowing that AI is transforming logistics, without knowing what to buy on Monday.

The sector's own statistics tell a different story. Transportation and storage is the second-least AI-adopting sector in the European Union, and AI applied to logistics is the least-used application of all.

So this guide answers three questions the category avoids. How many companies are doing this, and how many of those succeed. What the capabilities cost and what they get right, with the arithmetic shown. And what European law asks of a company that buys AI. Most of what is written about the AI Act is written for the people building it.

For a company that owns goods and hires carriers, AI in logistics means a small set of specific things: reading transport documents, predicting arrival times, forecasting volumes and suggesting routes, all running inside the systems that already handle the transport. It does not mean robots or autonomous trucks, and it is not one product you buy.

Chart comparing AI adoption in EU transportation and storage (11.2%) with the all-sector average (20.0%) in 2025

How Much AI Is Running in European Logistics Today?

Less than in any sector except construction. Eurostat's 2025 figures put 20.0% of EU enterprises with ten or more staff using AI technologies, up 6.5 percentage points from 13.5% in 2024. Transportation and storage sits at 11.2%, roughly half the all-sector average and second-lowest of every sector measured. Only construction is lower, at 10.8%.

Company size matters more than the headline suggests. Small enterprises with 10 to 49 staff are at 17.0%, medium enterprises with 50 to 249 at 30.4%, and large enterprises with 250 or more at 55.0%. A mid-market shipper sits across the medium band and the bottom of the large one, so the honest expectation is that somewhere between a third and a half of companies your size use AI for something.

For something is doing a lot of work in that sentence. Eurostat also asked what AI-using enterprises use it for, and among the purposes it measures, AI in transportation and logistics came last, at 6.08%. Marketing and sales reached 34.7%. Business administration reached 31.05%. Read that carefully, because the base is enterprises already using AI rather than all enterprises. Even so, the ranking carries the point. Among companies that have adopted AI, almost none have pointed it at freight.

Geography compounds it. Denmark leads at 42.0%, Finland at 37.8%, Sweden at 35.0%. Romania is lowest in the European Union at 5.2%, with Poland at 8.4% and Bulgaria at 8.5%. A Romanian shipper looking at AI is operating in the least-adopted sector in the least-adopted country in the bloc.

Adoption is also not the same as working. BCG surveyed more than 180 logistics providers and shippers in January 2026 and found about 10% had scaled AI across core operations. In Europe that figure was 6%. Only 13% reported measurable value from embedding AI into daily operations, in things like unit costs, service levels or margins. On the shipper side specifically, 1% had embedded AI in core logistics processes and 7% could point to measurable improvements, while 69% were still exploring or piloting. The sample is small and weighted toward providers, so treat it as a direction of travel.

Worth holding one vendor number up against the rest, because you have probably seen it. Descartes' ninth annual transportation management benchmark study, published September 2025 with 616 shippers and logistics providers surveyed by SAPIO Research, reports that 96% are "using generative AI" in transportation management. The same release reports that only 17% have fully automated processes and that over a third remain heavily or mostly reliant on manual work. Both numbers are from the same document. The 96% measures whether anyone in the company has touched the technology. The 17% measures whether it changed how the work gets done. When a vendor survey tells you adoption is near-universal, that is usually the gap it is hiding.

The obstacles are consistent across the data. Among EU enterprises that considered AI and did not adopt it, roughly seven in ten cited a lack of in-house expertise. BCG found the same thing from the other direction. The barriers respondents named most were unclear return and internal capability, ahead of cost or technical difficulty.

If you want to measure where you stand before buying anything, the logistics KPI baseline is the place to start. [INTERNAL LINK NEEDED: logistics KPIs article]

Bar chart of EU AI adoption by sector in 2025, with transportation and storage at 11.2% and a split by company size: 17.0% small, 30.4% medium, 55.0% large

Which AI Capabilities Work in Road Freight Today?

Three have real published evidence, one widely-sold category has none at all, and one gets less space here than its prominence elsewhere would suggest. The benefits of AI in logistics are real and they are narrower than the category implies.

Document Extraction Has the Strongest Evidence, and a Ceiling

Reading a transport document and turning it into structured data is the most evidenced AI capability in road freight, because academic researchers have measured it on real freight paperwork.

Chen and colleagues, publishing in Electronics in May 2025, ran an OCR-plus-language-model pipeline over invoices, shipping documents and bills of lading. On the public SROIE benchmark they reached 95.5% average field precision and 91.5% whole-document extraction accuracy. On real invoices from a Taiwanese shipping company, field precision held at 97.15% while whole-document accuracy fell to 85.29%. Real freight paperwork is harder than benchmark paperwork.

Hu and colleagues, in Scientific Reports in August 2025, tested the same kind of extraction on non-standardised logistics documents, including customs declarations, logistics labels and bills of lading. They reported F1 scores of 80.9 to 83.9 on public datasets, with whole-document accuracy of 88.85% on one and 93.3% on the other, and a 5 to 8% accuracy advantage over other multimodal models on a private customs dataset. Worth reading the whole table: on one of the public sets, two of the baselines it was tested against slightly outperformed it.

The number that matters most comes from the DocILE competition at ICDAR and CLEF in 2023, run over 6,680 annotated real business documents. On key information localisation and extraction the winning system reached roughly 71% F1. On line item recognition it also reached about 71% F1. Line items are the goods section of a freight invoice or a CMR. The competition was co-organised by Rossum, a document-AI vendor, which is worth knowing, though a public academic competition with a published leaderboard is stronger evidence than any vendor's own figure.

Put those together and the picture is consistent. Individual fields come out right 95 to 97% of the time. Whole documents come out right 85 to 93% of the time. Line item tables come out right about 71% of the time. A review queue is part of the product. Anyone selling you unsupervised document processing for freight is selling past the evidence.

What ETA Prediction Actually Delivers

Machine learning improves arrival time prediction. A 2025 review in PeerJ Computer Science compiled the reported gains by technique:

  • Refining base maps with high-resolution data reduces prediction error by up to 20%.
  • Segmenting drivers into behavioural archetypes improves accuracy by up to 12%.
  • Integrating braking telemetry reduces error by 9 to 11%.
  • Hybrid algorithmic and human-heuristic methods improve accuracy by 15 to 20%.

Those figures come from the literature the review surveys rather than from one controlled experiment, so read them as a range of reported results.

The most-quoted figure in this area is DeepMind's, from work on Google Maps published in September 2020: up to a 50% reduction in ETA inaccuracies, with Taichung City at 51% and London and Copenhagen at 16%. It is nearly six years old and it measures consumer road ETA rather than multi-leg freight ETA, so it tells you the technique works and very little about your lanes.

One prerequisite runs through all of it. The same PeerJ review notes that without an accurate base map, even advanced models produce unreliable results. ETA quality is a data-quality question before it is a modelling question.

Demand Forecasting Has One Number, Recycled for a Decade

Everyone cites McKinsey's figure. Applying AI-driven forecasting to supply chain management can reduce errors by between 20 and 50%. It comes from a February 2022 article with no sample, no methodology and no independent verification, and the range traces back to McKinsey's own 2017 work, which means it predates large language models entirely. McKinsey also sells this work. Use it as an indication of the order of magnitude and nothing more.

Why Route Optimisation Gets One Paragraph

Route optimisation appears on 15 of the 17 pages currently ranking for this topic, more than any other use case, and it gets the least space here. The reason is the evidence. There is no published accuracy evidence for freight route optimisation comparable to the document sources above. Alpega names predictive routing and predictive ETA on its product pages with no mechanism described and no accuracy figures. FarEye publishes blog articles about ETA accuracy while publishing no ETA accuracy figure on any product page. The capability is real and widely deployed. What is missing is anything you could hold a vendor to.

How Far Along Are AI Agents in Freight?

Agentic AI in supply chain is the newest layer, and vendor claims are uneven. Transporeon states in one sentence that its Trimble Arc Agent is live today with all its listed skills. project44 lists ten agents as generally available. Sennder names four agents and states no availability status for any of them.

Gartner's view is worth holding alongside that. In June 2025 it predicted that over 40% of agentic AI projects will be cancelled by the end of 2027, citing escalating costs, unclear business value and inadequate risk controls. It also estimated that only around 130 of the thousands of vendors claiming agentic AI are genuine, a practice it calls agent washing. That is a forward-looking prediction rather than a measured failure rate, and it has not happened yet. It is still the right frame for reading an agent roadmap.

Freight Invoice Audit Has No Credible Figure

This one deserves naming plainly, because the numbers circulate widely. The claim that 5 to 10% of freight invoices contain billing errors appears on at least a dozen vendor blogs with no study, no sample, no author and no date behind any of them. The claim that one in four invoices gets rejected traces to a trade magazine with nothing underneath it. The claim that invoice errors cost 2 to 5% of total transportation costs is attributed only to "some sources". We could not find a primary source for any of them. If a vendor quotes you one of these, ask where it comes from.

What the Paperwork Costs, According to the Commission

The one official quantification of the problem is the European Commission's impact assessment for the electronic freight transport information regulation, published 17 May 2018. It estimated that 99% of cross-border transport operations in the EU still involve paper-based documents at one stage of the operation or another, which is a narrower claim than the way it usually gets quoted. It put the administrative burden at more than 380 million hours a year, worth almost EUR 7.9 billion, of which road transport carried EUR 5,962 million, about three quarters of the total. Its two case studies found time savings worth EUR 9 to 37 and EUR 21 to 87 per trip.

Those estimates are for 2018 and they are more than eight years old. They remain the only official EU figures on the subject, which is why everyone cites them, so always say the Commission's 2018 impact assessment estimated rather than the EU says.

Chart of AI document extraction accuracy: 95–97% per field, 85–93% per whole document and about 71% on line items, showing why human review is needed

What This Means for 300 Loads a Month

Nobody in this category shows the arithmetic, so here it is with every assumption labelled.

Take a shipper running 300 loads a month with three documents per load, which is a transport order, a CMR and a carrier invoice. That is 900 documents a month.

The prize first. The Commission's 2018 impact assessment put time savings from digitalising paper transport processes at EUR 9 to 37 per trip in one case study and EUR 21 to 87 in the other. At 300 loads a month, the first case study works out to roughly EUR 2,700 to 11,100 a month of savings available, and the second to EUR 6,300 to 26,100. Read that as the size of the prize rather than the size of today's bill. That multiplication is ours, applied to the Commission's per-trip range, and the underlying figures are from 2018.

Now the error side. At the 85.29% whole-document accuracy Chen and colleagues measured on real shipping-company invoices, 132 of those 900 documents come out wrong somewhere. At the cleaner 91.5% benchmark figure, 77 do. Assume two to four minutes to find and fix each one, which is our assumption rather than a measured figure, and 132 corrections is roughly four to nine hours a month of review work that has to exist and be staffed.

Two caveats. Chen's 85.29% was measured on invoices, so applying it to a mix of orders, CMRs and carrier invoices is an inference. And the CMR goods section is the harder case, which is what the DocILE line item figure at about 71% is telling you.

Substitute your own numbers. Loads a month, documents per load, and your own loaded hourly rate for whoever does the checking. The ratio between the prize and the review queue is the whole business case, and it is a calculation no vendor will do for you.

Run this calculation on your own lanes. Bring your monthly loads, your document mix and your hourly rate, and we will work through the prize and the review queue with you. Book a demo

What Accuracy Should You Require in Writing?

A figure with only one number attached cannot be compared to anything. An accuracy claim for a prediction needs both a tolerance window and a horizon, and most of the market publishes one or neither.

Shippeo publishes 90% accuracy for road delay predictions up to 12 hours prior to delivery. That is a horizon with no tolerance window, which makes it unfalsifiable, because accurate to within fifteen minutes and accurate to within four hours are both consistent with the sentence.

project44 publishes the only figure in the market with both. It reports a +28 percentage point improvement in accuracy at the 10-hour horizon, within a ±2-hour window, measured across more than 500 full-truckload shippers. Two things to note before you use it. It is an improvement delta with no stated baseline, so no absolute accuracy level can be derived from it. And it sits on an engineering blog rather than a product page, which makes it a published result rather than a commitment. The other ten vendors in our audit publish no ETA accuracy figure at all.

On document extraction the gap is total. We looked for a published document-extraction accuracy figure from twelve TMS, visibility and transport-document vendors selling into the European market in August 2026, across roughly fifty of their own pages, and did not find one. Deliwell included. Four of the twelve are not European companies, since project44 and e2open are American, Descartes is Canadian and FarEye is Indian, but they all sell here.

Be precise about what that claim means. It means none of them publishes such a figure publicly. It does not mean none of them has measured one. Five gaps bound it. project44's ETA datasheet sits behind a download gate. e2open's classification brief is gated. None of the twelve publishes public versioned release notes. Non-English locales were not systematically swept, which matters for a claim about vendors selling across Europe. And Dashdoc's AI help-centre index returns a 404, which matters because Dashdoc has the most developed document-extraction product in the set. Sales-only material sits outside all of it: RFP responses, service level agreements and security questionnaires are where extraction-accuracy commitments normally live in this market, and none of that is public.

So the question for a vendor is not whether their extraction is accurate. It is accuracy on your document mix, measured how, over what sample, and with what resulting review-queue rate. If you are evaluating AI logistics software, or an AI TMS, that is the line to put in the RFP.

The audit also turned up five distinct postures worth recognising when you read a vendor's AI page. Shippeo publishes accuracy levels and describes no current human oversight, only a future intention. project44 publishes both the model and the oversight, with agents running inside customer-defined guardrails and a transparent workflow history. Transporeon publishes the strongest oversight language in the whole audit, stating that every exception is passed to a person, alongside a procurement capability that runs 80% of spot volume autonomously. Dashdoc and Cargoson, the two mid-market products, publish the human safeguard and no model figures at all, with Dashdoc describing a review screen and Cargoson stating that its team and its AI digitise documents together and that humans remain in charge. One vendor, e2open, advertises publishing demand forecasts to planning systems without human review, though that is about forecasts rather than document checking.

Why Do Most Logistics AI Projects Not Reach Production?

The failures trace to data readiness and to people. Algorithms are rarely the problem, which means the disadvantages of AI in transportation are mostly not technical ones.

That follows directly from the accuracy evidence above. If a review queue is part of the product, then adopting document AI is a commitment of people, and every serious study of AI failure lands on data and staffing. Models come up much less often.

Gartner's most-quoted prediction is also the most-mangled. The full sentence is that through 2026, organisations will abandon 60% of AI projects unsupported by AI-ready data. The conditional is the statistic. It is not a prediction that 60% of all AI projects fail. Alongside it, Gartner found that 63% of organisations either do not have or are unsure whether they have the right data management practices for AI, from a survey of 1,203 data management leaders in July 2024.

Gartner's supply-chain-specific figure points at training. In May 2025 it predicted that by 2028, 60% of supply chain digital adoption efforts will fail to deliver promised value due to insufficient investment in learning and development, from a survey of 579 supply chain practitioners. The stated cause is not the technology.

Its April 2026 survey of 140 senior supply chain leaders at companies with $250 million or more in revenue found 56% citing integration with legacy systems as a major challenge and 50% citing limited internal expertise. That is a small panel of large companies, so it describes the enterprise end of the market rather than yours.

Deloitte's numbers are the most concrete on the pilot-to-production gap. Surveying 2,770 director-level and above respondents in May and June 2024, it found that 68% had moved 30% or fewer of their generative AI experiments into full production, and that 55% had encountered data obstacles that stopped them pursuing certain use cases.

McKinsey's 2026 State of AI, fielded in May and June 2026 with 1,719 respondents, found that 37% attribute at least some EBIT impact to AI, unchanged year on year. That leaves 63% who cannot. Only 6% qualify as high performers attributing 5% or more of EBIT to AI. McKinsey sells AI transformation work, which is worth knowing when reading its numbers.

RAND's 2024 study is the most useful for a practitioner, because it is qualitative and specific. From 65 interviews with data scientists and engineers, it identified five root causes. Three apply directly to a mid-market shipper: leadership failing to communicate to the technical team which problem needs solving, organisations lacking sufficient high-quality data to train performant models, and organisations applying AI to problems beyond the current state of the art. That paper also repeats a figure that more than 80% of AI projects fail, attributed to a magazine article rather than to any study. The number should not be attributed to RAND as a finding.

There is one failure mode specific to supply chain rather than to AI generally, and a 2025 systematic review of 66 studies in Information names it. Counterparties will not share data. Marketing AI runs on first-party data. Supply chain AI runs on data held by carriers, suppliers and customers who have no obligation to give it to you and sometimes a reason not to. That is why the document flow you already own is a better starting point than a prediction that depends on someone else's telemetry.

One number you have probably seen deserves taking apart. The claim that 95% of AI pilots fail comes from a July 2025 MIT NANDA paper, and it does not survive inspection. Around 80% of the companies in its denominator never piloted any custom generative AI at all, so the 5% success rate is measured against a population that mostly never tried. The bar for success was a marked and sustained profit-and-loss impact within six months, which fails any project that merely broke even. The sample was 52 executive interviews and 153 survey responses, it was labelled preliminary and was not peer reviewed, and all four authors work on the agentic AI frameworks the paper recommends. A detailed critique on the 80,000 Hours podcast in April 2026 walked through each of those problems. The honest replacement is McKinsey's 63% with no measurable EBIT impact, which at least states what it measures.

Chart of four AI failure and adoption figures from Gartner, Deloitte and McKinsey, 2024–2026, each labelled with what it measures and its sample size

What Does AI in Logistics Cost a Mid-Market Shipper?

The published figures are enterprise-shaped, and the fixed cost does not scale down with your headcount.

Gartner's cost model from February 2024 is the most detailed public estimate. For a document search and summarisation use case it puts upfront cost at $750,000 to $1 million and recurring cost at $790 to $1,200 per user per year, assuming 1,000 users and a six-month deployment. Coding assistants come in at $100,000 to $200,000 upfront for 200 users. A fine-tuned large language model for a regulated industry runs $5 million to $6.5 million.

Three caveats before anyone quotes those at you. The model dates from February 2024, and inference costs have fallen substantially since, with Gartner itself predicting in March 2026 that inference on a trillion-parameter model will cost providers over 90% less by 2030 than in 2025. The figures are Gartner's modelled estimates rather than observed spend. And the user-count assumptions are enterprise-scale.

That last one is the problem for a mid-market shipper, and it is worth stating plainly. A company running 50 to 5,000 transports a month might have five to thirty people touching a transport system. The upfront engineering, integration and governance cost does not fall proportionally as the seat count falls, so per-user economics get worse the smaller the operation. Spreading a $750,000 build across fifteen users is a different business than spreading it across a thousand.

Which is the argument for buying a capability that already sits inside a system doing the work. The alternative is running an AI project. The build cost is somebody else's, amortised across their customer base, and the integration work is already done because the data is already in the platform.

Worth noting what most vendors do publish, since it is rarely model performance. Transporeon publishes a 12% cost saving and 84% automation rate. Alpega publishes a 150% return on investment. Dashdoc publishes up to three times faster invoice processing, which is a speed claim rather than an accuracy one. Outside project44 and Shippeo, the published numbers in this market are commercial outcomes, and a commercial outcome does not tell you whether the model is any good.

What Does the EU AI Act Require of a Mid-Market Shipper?

Three articles apply to you today. The high-risk obligations that most compliance content warns about apply from 2 December 2027, and the ones for AI built into products from 2 August 2028.

Those dates moved recently and a lot of published guidance is now wrong. Regulation (EU) 2026/1744, the Digital Omnibus on AI, dated 8 July 2026, was published in the Official Journal on 24 July 2026 and entered into force on 27 July 2026. It amended the AI Act and pushed the Annex III high-risk obligations from 2 August 2026 to 2 December 2027. The AI Act itself still became generally applicable on 2 August 2026, so nothing was cancelled and nothing was delayed wholesale. One section moved. Anything you read describing an August 2026 high-risk deadline predates that regulation, and the Commission's own AI Act Service Desk still serves pre-amendment article text with a disclaimer saying so.

Articles 4, 5 and 50 bind you today. All three are covered below. Article 50 is the most recent of them, generally applicable since 2 August 2026.

Article 4, AI literacy, has applied since 2 February 2025. It requires providers and deployers to take measures to support the development of AI literacy among their staff, taking account of their technical knowledge and the context of use. It binds every deployer with no risk threshold and no exemption for small companies, so it reaches you the moment anyone in your business uses a bought AI feature. The Omnibus added a sentence clarifying that it does not require guaranteeing any specific level of literacy in any individual. Article 4 is absent from the list of provisions in Article 99(4) that carry AI Act fines, so claims that breaching it exposes you to a EUR 15 million penalty are not supported. Any penalty would be national, and that is untested.

Article 5 prohibits certain practices outright, and the penalty tier is the highest in the regulation. Fines run up to EUR 35 million or 7% of total worldwide annual turnover, whichever is higher, with small and medium enterprises paying whichever is lower. One prohibition matters for road freight. Article 5(1)(f) bans AI systems used to infer emotions of a natural person in the workplace, excepting systems intended for medical or safety reasons.

The line runs through the data flow. The sensor is not what decides it. An in-cab camera detecting drowsiness and alerting the driver, in the cab, without retaining the output, sits inside the safety exception. The same camera scoring frustration or attitude, with the score written to a fleet dashboard and used in coaching, bonuses or discipline, does not. Buying the system does not insulate you, because Article 5 prohibits its use. This is unsettled ground. There is no case law and no guidance on where fatigue detection ends and prohibited emotion inference begins, and it is arguable whether biometric fatigue signals count as emotions at all. Also worth knowing is Article 5(1)(g), which prohibits biometric categorisation inferring protected characteristics. In road transport the live one is trade union membership.

Article 50, transparency, has applied since 2 August 2026. Where an AI system interacts with people they have to be told, and synthetic output has to be marked. For a shipper that reaches a customer-facing tracking chatbot and any generated document. The grace period for marking synthetic output from systems already on the market expires 2 December 2026. Unlike Article 4, Article 50 is in Article 99(4)'s fineable list.

What Is Not High-Risk, and What Is

Most of what a shipper would buy falls outside the high-risk regime, and it is worth understanding why.

ETA prediction, route optimisation, demand forecasting and transport-document extraction are not high-risk. The relevant entry is Annex III point 2, which covers AI systems intended to be used as safety components in the management and operation of critical digital infrastructure, road traffic, or the supply of water, gas, heating or electricity. Its conditions are cumulative: the system has to be a safety component and it has to be used in managing or operating road traffic. Recital 55 narrows safety components further, to systems used to directly protect the physical integrity of critical infrastructure or the health and safety of persons and property. Road traffic in this context means the road traffic system itself, meaning signal control, traffic management centres, tunnel and motorway control. A shipper managing its own consignments and its own hired trucks is doing neither thing. That reading is ours rather than a regulator's, and you should form your own with your own counsel, but the cumulative structure of point 2 is on the face of the text.

Two things a shipper might buy are high-risk. AI recruitment screening falls under Annex III point 4(a), which covers recruitment or selection of natural persons including analysing and filtering job applications. That is the clearest case in the regulation. And AI-assisted tour allocation falls under point 4(b), which covers allocating tasks based on individual behaviour or personal traits and monitoring or evaluating the performance and behaviour of persons. Allocation based only on vehicle capacity, licence class, geography and driving-time constraints, with no individual behaviour or traits involved, is arguably outside 4(b), though that is untested and depends on the specifics.

Carrier performance scoring usually sits outside the AI Act, because point 4 covers workers and carriers are business counterparties. There is an exception worth knowing, and it is the most under-appreciated exposure in the whole area. Point 4 extends to access to self-employment, so where the carriers being scored are owner-drivers or sole traders, 4(b) may engage. In European road freight that is a substantial share of the carrier base. It is also untested.

There is a derogation, and it will not rescue the interesting cases. Article 6(3) lets an Annex III system escape high-risk classification where it does not pose a significant risk of harm, and the qualifying conditions are alternatives rather than cumulative. But the same article states that a system shall always be considered high-risk where it performs profiling of natural persons, which covers driver behaviour scoring, candidate ranking and behaviour-based allocation. Articles 6(4) and 49(2) place the documentation and registration duties on the provider, and Article 49(3) limits deployer registration to public authorities. A private shipper never registers anything. Your obligation is a procurement one, which is to ask the vendor for its Article 6(3) documentation.

Shippers are widely told they must run a fundamental rights impact assessment. They do not have to. Article 27 reaches bodies governed by public law, private entities providing public services, and deployers of the creditworthiness and insurance systems in Annex III points 5(b) and 5(c). It expressly carves out Annex III point 2. A private commercial shipper is none of those, even for recruitment systems. You do still owe a GDPR data protection impact assessment, which is a different instrument.

The trap that does apply is Article 25(1)(c). A deployer becomes a provider, with the full obligation set, where it modifies the intended purpose of an AI system such that the system becomes high-risk. Point a general-purpose scoring or ranking tool at job applicants or at owner-drivers and you may have converted yourself into a provider, inheriting risk management, data governance, technical documentation, logging, human-oversight design, accuracy requirements, a quality management system, conformity assessment, CE marking and registration. Fine exposure moves with it. How much configuration counts as substantial modification has no threshold in the text and no guidance behind it.

From 2 December 2027, deploying a high-risk system brings a set of duties under Article 26. You use it per the provider's instructions. You assign human oversight to people with the competence and the authority to exercise it. You monitor operation and suspend use if a risk appears, and you keep the logs for at least six months.

One of those duties has a longer lead time than the rest. You must inform workers' representatives and affected workers before putting the system into service. In unionised European road freight that is a scheduling problem. You cannot do it the week you go live. Article 26(9) is the hinge into the next section: you use the information the provider gives you to carry out your GDPR impact assessment.

Two things to check before relying on any of this. The Commission's guidelines on high-risk classification were due no later than 2 February 2026 and, as of writing, exist only in draft. DLA Piper's read of that draft, published 19 May 2026, is that it does not address Annex III point 2 at all. And the amended Article 113, which carries the new dates, sits in Article 1(23) of Regulation 2026/1744 and should be read in the Official Journal before anyone quotes its exact wording.

Matrix of which EU AI Act obligations apply to a road-freight shipper and from when, splitting high-risk from non-high-risk systems

What About Driver GPS, Carrier Scores and GDPR?

GDPR is what fines shippers today. Every enforcement decision we found turned on three things: sampling frequency, retention period and transparency.

Employee consent will not carry driver tracking. The European Data Protection Board's 2020 consent guidelines note that an imbalance of power also occurs in the employment context. The earlier Article 29 Working Party opinion on data processing at work is blunter, finding that employees are almost never in a position to freely give, refuse or revoke consent. The bases that do work are legitimate interests under Article 6(1)(f), with a documented balancing test, legal obligation under Article 6(1)(c) for tachograph and driving-time records under Regulation (EU) No 165/2014 and Regulation (EC) No 561/2006, and contract performance under Article 6(1)(b) for narrow purposes.

A data protection impact assessment is required in practice. Article 35(1) requires one where processing is likely to result in high risk, and Article 35(3)(a) makes it mandatory for systematic and extensive automated evaluation of personal aspects on which decisions are based. Article 35(4) has each national authority publish its own list of processing that requires one, and systematic employee monitoring appears on most of them, so check your own market.

One control the Working Party guidance gives you is worth implementing regardless. An employee should in principle be able to turn location tracking off temporarily where circumstances justify it, such as a medical appointment, and employees must be told a tracking device is installed in a vehicle they drive.

The human-in-the-loop defence is weaker than most people think. In SCHUFA Holding, decided 7 December 2023, the Court of Justice looked at automated probability values. It held that establishing one amounts to automated individual decision-making under Article 22, where a third party receiving that value draws strongly on it to establish, implement or terminate a contractual relationship. So if a score is what the decision-maker relies on, the score is the decision, and a rubber-stamp approval does not change that.

Applied to freight, that has two consequences. Removing a driver from tour allocation on the strength of a behaviour score sits squarely inside Article 22(1), because loss of work is at least a similarly significant effect. Carrier de-listing is a two-step question: de-listing a limited company falls outside GDPR entirely, because GDPR protects natural persons, but de-listing an owner-driver or sole trader engages it fully. SCHUFA was a consumer credit case and no court has applied it to freight scoring, so treat it as a strong analogy rather than settled law.

There is no single European answer on driver monitoring, and Article 88 is why. It lets member states set more specific employment rules by law or collective agreement, expressly including workplace monitoring. Romania's Law 190/2018 is the sharpest example for this audience. Monitoring by electronic means is lawful there only where five conditions hold together: the employer's legitimate interests prevail, employees were informed in advance, the employer consulted the trade union or employee representatives beforehand, less intrusive means would not achieve the purpose, and retention does not exceed 30 days except in justified cases. Neither the consultation requirement nor the 30-day cap has an EU-level equivalent. Those five conditions are as rendered by Eurofound, which is the accessible source, so confirm against Monitorul Oficial before relying on them. Italy takes a different route, requiring a collective agreement or Labour Inspectorate authorisation under Article 4 of the Workers' Statute, and the Garante has held that a collective agreement is necessary without being sufficient.

The enforcement record is small and consistent, and all four decisions below are as reported in secondary sources rather than read in the decision texts. The Italian Garante fined a road transport company EUR 6,000 in 2026 over a GPS system recording vehicle position every 60 seconds with real-time monitoring, which it characterised as essentially continuous tracking of employees. A collective agreement was in place and did not save it, and the fine was reduced after the company moved to 15-minute intervals and disabled the real-time view. In March 2025 the same authority fined a transport company EUR 50,000 over roughly 50 employees tracked without proper information, with data retained beyond 180 days and required vehicle stickers missing. Romania's ANSPDCP fined UP România EUR 4,000 in November 2024 for processing an employee's GPS location during leave, with retention beyond the national 30-day limit. France's CNIL fined a moped rental operator EUR 125,000 in 2023 over GPS collected every 30 seconds, though those were customers rather than employees, so it is useful only for the principle about sampling frequency.

In the employee-monitoring cases, fines ran EUR 4,000 to EUR 50,000. Which means the real cost of getting this wrong is works-council friction, litigation and the reputational damage of a published decision. The fine is the small part.

What Should You Not Automate?

Four things, and all four share one property: the cost of a silent error is paid by someone who cannot see it.

Any output that is a decision about a person, unless a named human can override it and does. The test is evidence. Can you show a case where the human overrode the system, and what happened next? A review step nobody has ever used is not oversight, and after SCHUFA it is not a defence either.

Any line-item extraction that posts straight to an invoice or a customs declaration. The test is a number you should already have: what is your measured review-queue rate on goods lines? If you cannot answer, the roughly 71% line-item accuracy ceiling is running unmonitored inside your billing.

Anything a works council would have to be told about after go-live rather than before. The test is a signature and a date. Article 26(7) makes prior notification an obligation from December 2027, and Romanian law requires prior consultation today, so the sequencing is the compliance question.

Anything whose input data belongs to a counterparty with no obligation to give it to you. The test is whether you have that data today, machine-readable, without asking anyone. If a capability depends on carrier telemetry you do not control, you are buying a dependency. This is the failure mode specific to supply chain AI.

One design point worth naming, because it answers the pattern in all four enforcement decisions above: driver GPS in Deliwell is scoped to active shipments rather than running as continuous personal monitoring. Sampling frequency and retention were what every one of those fines turned on.

Where Should a Mid-Market Shipper Start?

With the document flow you already own. It is the only capability in this guide with published accuracy evidence, and the only one whose input data you control.

The use of AI in logistics that pays for itself at this scale is narrow, and the sequence matters more than the tool.

  1. Pick one document type you handle every day. A CMR, a carrier invoice, a delivery note. One.
  2. Measure the baseline before touching AI. Minutes per document and current error rate. Without this you cannot tell later whether anything improved, and the studies above are unanimous that unclear return is what kills these projects.
  3. Require an accuracy figure on your own document mix. A benchmark number will not tell you anything about your paperwork. Ask for the sample size and the measurement method too.
  4. Keep a review queue and measure its rate. That number is your early warning system, and it is also the thing you renegotiate on.
  5. Only then consider prediction. ETA and forecasting depend on data you partly do not own, which is a harder project with a softer evidence base.

That order works because the data is first-party, the accuracy evidence exists, the failures show up where you can see them, and none of it touches anyone's employment.

The thing worth being clear-eyed about is what you are replacing. For most mid-market shippers the incumbent is not software at all. It is a shared mailbox, a folder of scanned CMRs, a spreadsheet tracking which ones came back, a monthly ERP export nobody reconciles, and a WhatsApp group for chasing the rest. An AI project built on top of that inherits every one of its data problems, which is RAND's second root cause almost word for word.

It is also worth noticing who touches those documents. Sales promised the delivery date, purchasing ordered the goods, the warehouse packed them, logistics booked the carrier, finance pays the invoice, and management is asked why transport cost moved. Six departments, one transport, and usually six different versions of where the paperwork is. That gap between the people is the problem AI gets pointed at, and it is not really a modelling problem.

Where Deliwell Fits

Deliwell is a road-freight TMS sized for mid-market European shippers, covering planning, allocation, dispatch, tracking, documents, compliance, invoice validation and KPIs. Proof of delivery and document handling sit inside it as stages of the transport rather than as separate products, so the consignment note comes off a transport order that is already in the system. [INTERNAL LINK NEEDED: proof of delivery software article]

On the AI specifically, described by what each thing does:

  • Extraction from PDF, Excel, images and forwarded email, with human validation before submission. Which is the design the accuracy evidence above argues for.
  • Smart document generation from data already in the transport order, so the consignment note is not retyped.
  • Predictive ETA and anomaly detection.
  • Route optimisation suggestions, which propose a sequence and leave the dispatcher to accept or change it.
  • Status-triggered notifications and compliance automation, including e-Transport/UIT.

The exact behaviour behind the last three is worth asking about in a demo, and it is worth asking any vendor the same question. What the model reads, what it emits, and who checks it.

Three capabilities are on the roadmap rather than in the product, and they are worth naming because two of them are what most AI-in-logistics content promises. Carrier performance scoring, demand-based planning and automated exception resolution are not live. Carrier scoring is also the capability carrying the most regulatory weight, for the reasons in the AI Act and GDPR sections above, which is a good reason to ask any vendor selling it exactly how their scoring handles owner-drivers.

Deliwell is in production with mid-market shippers including Stihl, Sonepar and Frigoglass. Standard go-live is measured in days, and the first transport request usually goes out within 30 minutes. ERP integration is optional, and where it is in scope it adds time.

Where it fits badly is worth saying in the same breath. Under 50 shipments a month, the platform is more than the operation needs. Five or more GPS providers in the carrier base makes the integration work heavier. Complex air and sea costing is supported on an all-in basis only, because Deliwell is road-first. Specialised freight is a poorer fit, and special transport for heavy, oversized and hazmat loads is a paid add-on module.

Deliwell pricing is fixed, starting at €250/month and scaling with your transport volume — delivered as an all-in-one SaaS subscription with no hidden IT costs. Contact Deliwell for a pricing proposal tailored to your operation. Nobody in this market publishes a price. That is why ours is on this page.

On the same principle, there are cross-customer benchmarks available for manual-work reduction and finance reconciliation time that are not in this article. Putting an unattributed percentage here would undercut everything above it about unattributed vendor percentages.

Deliwell can run a proof of concept on your own lanes, your own carriers and your own paperwork, long enough to measure a real review-queue rate. Book a demo

FAQ

What is AI in logistics?

For a company that owns goods and hires carriers, it means a specific set of capabilities running inside existing transport systems: reading and generating transport documents, predicting arrival times, forecasting volumes and suggesting routes. It does not mean warehouse robots or autonomous trucks, and it is not a single product. Eurostat's 2025 data puts AI use in transportation and storage at 11.2%, the second-lowest of any EU sector.

Is AI in logistics worth it for a company running 300 loads a month?

It depends on a calculation you can run yourself. At 300 loads and three documents each, that is 900 documents a month. The Commission's 2018 impact assessment implies roughly EUR 2,700 to 11,100 a month of available savings at that volume in the first of its two case studies. Against that, published extraction accuracy of about 85% on real shipping invoices means around 132 documents a month need correcting, which is four to nine hours of review work. Compare the two using your own hourly rate.

Which AI use cases in logistics have published accuracy figures?

Very few. Document extraction has peer-reviewed figures: 95 to 97% at field level, 85 to 93% for whole documents and about 71% for line items. ETA prediction has published improvement ranges in academic reviews. Demand forecasting has one consultancy range from 2022. Among vendors, only project44 publishes an ETA figure with both a tolerance window and a horizon, and no vendor we checked publishes a document-extraction accuracy figure publicly.

Is driver GPS tracking legal in the EU?

Yes, with conditions, and the conditions are national. Consent from an employee is generally not a valid basis, so most tracking relies on legitimate interests or a legal obligation such as driving-time records. A data protection impact assessment is required in practice. GDPR Article 88 lets each member state add rules: Romania caps retention at 30 days and requires prior consultation with employee representatives. Enforcement decisions have turned on sampling frequency, retention and whether employees were told.

Can I use AI to score my carriers?

Legally, carrier scoring usually sits outside the EU AI Act's high-risk regime, because Annex III point 4 covers workers and carriers are business counterparties. Two qualifications matter. Where the carriers are owner-drivers or sole traders, point 4 extends to access to self-employment and may apply. And where an automated score drives de-listing of a natural person, GDPR Article 22 and the SCHUFA judgment apply regardless of the AI Act.

Case studies

Check out what our customers are saying about us!
Sonepar Romania electrical materials distribution warehouse
Sonepar logo

Sonepar Romania centralises transport reporting and cost control with Deliwell

How a 15-location electrical distributor replaced Excel, email, and WhatsApp with one platform — and gained branch-level cost visibility without adding IT complexity.
Melinda S.
Logistics Manager & Stock Controller, Sonepar Romania
Case Study
Inteva Products automotive components manufacturing facility
Inteva Products logo

Inteva gains full transport visibility and eliminates manual work with Deliwell

How Inteva replaced scattered spreadsheets, emails, and WhatsApp threads with one source of truth for every shipment, cost, and invoice.
Adriana H.
PC&L Manager, Inteva Closure
Case Study
Hamilton Central Europe laboratory equipment manufacturing facility
Hamilton logo

Hamilton Central Europe reduced transports costs by 15% by replacing their old TMS with Deliwell

How Hamilton replaced four disconnected tools, went live with UIT compliance in days, and cut shipping workload by 15%.
Roxana C.
Supply Chain Manager - Hamilton Central Europe
Case Study