An Accuracy Acceptance Test for a Roadway Asset Inventory
Define accuracy, sampling, and tolerance upfront—before vendors bid—or disputes will bury you.
Staff Writer · · 7 min read

An accuracy acceptance test for a roadway sign and asset inventory has one job: tell you, with a number you can defend in a contract dispute, whether the delivered dataset matches the road. Most agencies skip the hard part and write "95% accuracy required" into the scope of work without defining what accuracy means, what gets counted, or how the sample gets picked. That single omission is where most inventory acceptance disputes start, and it's fixable before a single sign gets scanned.
The test has three moving parts: a ground-truth sampling frame that isn't secretly biased toward easy signs, a positional tolerance that matches what the data will actually be used for, and a plan for what happens when a corridor fails. Skip any one of these and the acceptance test becomes theater. The vendor passes, the agency signs off, and six months later a maintenance crew can't find a stop sign that the inventory swears is on the northwest corner of an intersection it's actually thirty feet south of.
## Building a ground-truth sampling frame that isn't rigged
Ground truth means an independent count and location of every asset in a set of test corridors, collected by someone other than the vendor, using a method more precise than what the vendor used. If the vendor drove mobile LiDAR at 30 mph, your ground truth crew needs a survey-grade GPS unit or a robotic total station, walking the corridor on foot. Validating a dataset with a method of equal or lesser precision than the one that produced it amounts to a coin flip, not a check.
The sampling frame is where most agencies get sloppy. If you let the vendor nominate the test corridors, you'll get their best three miles: flat terrain, wide shoulders, good sign retroreflectivity, minimal vegetation overgrowth. That tells you nothing about how the system performs on the rural two-lane with an 8-degree curve and sign faces buried in kudzu that make up the other 40% of your network.
A defensible frame stratifies by the conditions that actually predict detection failure: functional class (interstate, arterial, collector, local), posted speed, roadside vegetation density, and sign density per mile. Pull your corridor list before the vendor bids, seal it, and don't let anyone edit it once contract signature happens. Randomly select within each stratum, not just from the strata themselves, or you'll end up back at "the easy three miles" with extra paperwork attached.
Sample size matters more than most RFPs acknowledge. A ground truth set of 200 signs sounds like a lot until you remember that regulatory signs, warning signs, and guide signs each need their own accuracy figure, because a vendor's neural network trained heavily on stop signs and yield signs will often miss object markers and chevrons at a materially higher rate. Splitting 200 signs across a dozen categories leaves some categories with sample sizes too small to support any statistical claim. Budget for at least 30 to 50 instances per major sign category if you want a confidence interval that survives scrutiny, and say so in the RFP so the vendor isn't surprised by the math later.
## Setting positional tolerance around the use case, not around convenience
Positional tolerance is the horizontal distance between where the inventory says a sign is and where a survey-grade measurement says it actually is. There is no universal right answer here, and any consultant who quotes you a single tolerance figure without asking what you're using the data for is guessing.
If the inventory feeds a maintenance work order system where a crew drives to a GPS coordinate with a bucket truck, a tolerance in the 3 to 5 meter range is often workable, since the crew can see the sign once they're in the right block. If the inventory feeds an ADAS-adjacent application, or a digital twin meant to support connected vehicle infrastructure, you're in sub-meter territory, and the vendor's collection method (mobile LiDAR vs. mobile imagery with photogrammetry vs. handheld GPS) needs to match that bar before you even get to testing.
MUTCD compliance reviews and retroreflectivity management programs sit in between: you need to find the right sign fast enough to inspect it, but you're not driving a robot to the coordinate. A 2 to 3 meter tolerance usually clears that bar. Whatever number you choose, write it into the contract as a pass/fail threshold per asset, not as an average. A dataset with an average positional error of one meter can still contain a subset of assets that are consistently 15 meters off because of a georeferencing error in one flight strip, and averages will hide that pattern from you.
## The recall number nobody wants to talk about
Positional accuracy only matters for signs the vendor actually found. Recall, the percentage of real-world assets that show up in the delivered inventory at all, is the metric that decides whether the dataset is usable, and it's the one vendors have the strongest incentive to leave loosely defined.
A vendor can hit 98% positional accuracy on the signs it detected while missing 15% of the signs that exist in a corridor, and a poorly written acceptance test will still call that a pass, because nobody wrote a recall floor into the contract. Recall failures cluster. They don't spread evenly across your network; they show up in the corridor with the sharp curve where the LiDAR unit's field of view swept past a sign at an angle, or in the stretch of road where overgrown vegetation occluded half the sign faces from the collection vehicle's sensors, or where a sign's retroreflective sheeting had degraded enough that a computer vision model trained on well-maintained signs didn't register it as a sign at all.
Testing recall at the network level and reporting a single aggregate number therefore produces a misleading picture. A network-wide recall of 92% could mean the vendor missed 8% of signs evenly, which is a manageable data quality issue, or it could mean the vendor missed 40% of signs in three specific corridors and near-zero elsewhere, which is a sensor or algorithm failure that will repeat itself on every corridor sharing those conditions. Telling the two apart requires a stratified ground truth frame from the outset, which is the other reason the sampling design matters as much as the pass/fail threshold.
## When recall fails in a single corridor
Here is the scenario that acceptance testing has to be built to catch: overall network accuracy clears the contract threshold, but one test corridor comes in badly under recall, say 70% against a required 90%, while every other corridor passes comfortably. The vendor will, reasonably, argue that the aggregate number should govern acceptance and that one bad corridor is noise.
Treat that argument with caution and investigate before ruling on it either way. A single-corridor failure is diagnostic information, not statistical noise, and the acceptance test needs a defined process for what happens next rather than a judgment call made under deadline pressure. Pull the raw sensor data for that corridor before anything else. Was there a documented equipment fault, a lens obstruction, a GPS dropout from tree canopy, something the vendor's field log actually recorded? If so, that's a data collection issue with a fix: recollect the corridor and retest it in isolation, holding the rest of the delivery to its already-passing status.
If the raw data looks clean and the miss is a detection failure, meaning the sensor saw the signs and the algorithm still didn't classify them, that's a different and more serious problem. It suggests the vendor's model has a blind spot correlated with something in that corridor: sign age, mounting height, a nonstandard bracket, faded sheeting, an unusual sign shape. Averaging that corridor's failure into the network score and moving on would bury the signal. The better move is to ask whether that same condition exists anywhere else in your network that wasn't in the test sample, because if it does, the problem isn't one bad corridor. It's one instance of a systemic gap that the sampling frame happened to catch, and there could be more.
Write a contract clause in advance that requires root-cause documentation for any corridor-level recall failure past a set delta from the network average, rather than negotiating what "failure" means after the fact under change-order pressure. Reserve the right to expand the sample in any stratum where a failure surfaces, at the vendor's cost, before final acceptance. Vendors who are confident in their process won't object to this; the ones who push back hardest on a corridor-level retest clause are telling you something about how confident they actually are in the delivery.
## The number the contract actually needs
None of this works as a single line item. "95% accuracy" without a defined denominator, a stratified sample, a positional tolerance tied to use case, and a documented process for corridor-level failure amounts to a hope rather than a specification. Write the sampling frame into the RFP before bids go out, agree on tolerance with the end users of the data (maintenance crews, traffic engineers, GIS staff) rather than with the vendor, and build the corridor-failure clause into the contract before you need it, not after a curve on Route 9 comes back at 70% recall and everyone's staring at each other in a conference room trying to figure out whose problem that is.
More in Features
The Privacy Schedule in a Street-Imagery Data Agreement
Gideon Lachance
Turning Repeat Passes Into a Change-Detection Layer
Nathaniel Kessler
Calibrating Low-Cost Vehicle Sensors Against a Reference Monitor
Luz Maribel Cervantes
