Finding a part you can't see: building the Spatial Maintenance Copilot
Spatial Maintenance Copilot is my entry for the OpenCV AI Competition 2026. It locates engine components hidden from view, and it stops only when its measured uncertainty is smaller than the part itself.

Table of contents
In brief
Key takeaways
- Spatial Maintenance Copilot locates engine components hidden from view by estimating the camera's pose and projecting each component's known 3D position into the image.
- Pose estimation and projection run deterministically in OpenCV 5. The LLM agent only chooses the next tool call and guides the user.
- The system stops when the uncertainty region around a predicted position is smaller than the component itself.
- Uncertainty starts from camera calibration, which reports a standard deviation for every parameter, and will be propagated to the component's position with Monte Carlo sampling.
Spatial Maintenance Copilot
My proposal for the OpenCV AI Competition 2026 was accepted for an AWS compute grant, so for the next few weeks I'm building the project in public. This is the first post in a short series about it.
The problem
The manual says the sensor sits behind the intake. You open the hood and see hoses, covers and brackets, but no sensor. So you guess, and a wrong guess costs time, or a part removed for nothing.
You know roughly where the component should be. You just can't confirm it with your eyes.
Geometry decides, the agent guides
You point a phone camera at the engine bay. The system estimates the camera's pose relative to the vehicle, projects the hidden component's known 3D position into the image and shows you where it is.
Two decisions shape the design:
- The geometry is deterministic. Pose estimation and projection run in OpenCV 5 and produce numbers you can check. The LLM agent never guesses where a part is. It decides which tool to call next and how to guide the user, based on what the geometry reports.
- Every estimate carries its uncertainty. A position without an error bar is just a confident guess.
For the demo I chose an engine bay over exterior body panels. Its depth and varied surfaces give pose estimation better conditioned geometry than large, nearly flat panels.
Knowing when to stop
The question I find most interesting is when the system should say "it's here".
Many vision demos stop when a model reports high confidence. I wanted a stopping rule tied to the physical world: the system stops when the uncertainty region around the predicted position is smaller than the component itself. Until then, the copilot asks you to move the camera to a viewpoint that shrinks that region.
That uncertainty comes from measured sources. Camera calibration already reports a standard deviation for every parameter, and the next step is propagating those, with Monte Carlo sampling, to the component's predicted position.
What's built so far
- A printed ArUco/ChArUco rig: a calibration board plus two independent reference boards that give the camera its pose in the engine bay. Board dimensions are measured after printing, not taken from the file, because a scale error in the board becomes a position error in the result.
- A camera calibration pipeline in Python and OpenCV 5 that keeps the uncertainty of each parameter, not only the parameter itself.
- A public monorepo under Apache 2.0, with a FastAPI perception service and a Next.js frontend, built to run on AWS.
What's next
- A ground truth atlas of component positions, triangulated from multiple views, so I don't have to take the engine apart to know where things are.
- End-to-end uncertainty propagation, validated against that atlas.
- A small user study before the submission deadline on 26 October.
The open question I'm most curious about: will the uncertainty estimates hold up on real footage, or turn out too optimistic? I'll write about the answer either way.
Follow along
The code is public from day one: Spatial-Maintenance-Copilot on GitHub. If you work in maintenance, robotics or computer vision and see a flaw in this approach, I'd like to hear about it.
Common questions
Why demo in an engine bay rather than on exterior body panels?
An engine bay has depth and varied surfaces, which give pose estimation better conditioned geometry. Large, nearly flat body panels make the camera's pose harder to pin down.
Why are the printed boards measured after printing?
A scale error in the board becomes a position error in the result. Board dimensions are measured on the printed boards instead of being taken from the design file.
How will the project check its own accuracy?
With a ground truth atlas of component positions, triangulated from multiple views, so the engine doesn't have to be taken apart. The uncertainty estimates are then validated against that atlas.