Different classes, cameras, or a private set?
The cone checkpoints prove the recipe. Your defect taxonomy, corridor, or camera network is a labeling campaign away.
Book a technical callDetection when you can name what you're looking for. Segmentation when you can't, because the whole point is catching what nobody predicted. The builds below are demonstrations, not the catalog.
The capability is adding your classes to a proven detector without losing the ones it ships with. The public proof is YOLOX extended with traffic cones, published in four sizes from an edge-camera S to a server-grade X. Weld defects, missing parts, or PPE follow the exact same run.
See the full benchmark →The capability is learning what normal looks like and flagging whatever breaks it. The public pilot watches motorways, where a load slides off a flatbed and gets flagged seconds after it lands. The same pattern guards any corridor, a conveyor belt, a rail track, a runway.
See the pipeline →Where teams point these capabilities
Here is the first capability with numbers attached. Four published checkpoints extend YOLOX with a traffic-cone class while keeping COCO, so you can see exactly what adding a class costs. Pick the size that matches your hardware.
| Checkpoint | Params | COCO mAP (kept) | Cone AP | Where it fits | Weights |
|---|---|---|---|---|---|
| yolox-pylon-s | 9.0M | 41.6 | 74.5 | Edge cameras and embedded boards, real time on a Jetson | Hugging Face → |
| yolox-pylon-m | 25.3M | 46.8 | 77.5 | The accuracy/speed sweet spot for most fixed cameras | Hugging Face → |
| yolox-pylon-l | 54.2M | 48.5 | 78.6 | Server-side inference over multiple streams at 3.7 ms per frame on an A100 | Hugging Face → |
| yolox-pylon-xl | 99.1M | 50.0 | 78.8 | Maximum accuracy for offline analysis and auto-labeling, 6.0 ms per frame on an A100 | Hugging Face → |
COCO mAP is measured after cone adaptation and shows the retained score on the original 80 classes, within 1.2 points of the official YOLOX val2017 baseline at every size and above it at s. Cone AP is at IoU 0.50 to 0.95 on the held-out cone split. See the full benchmark →
Cones are the demo class. The same run adds your classes, whether weld defects, missing parts, PPE, or whatever your cameras need to see.
Add your classesThe second capability, running as a pilot on motorway footage. No training set can enumerate everything a road will throw at a camera, so the pipeline learns the road itself and treats anything that interrupts it as a detection. The rarer and stranger the hazard, the more clearly it stands out.
A frame from the pilot, captured on the M25. Stage 1 has painted the drivable corridor blue, cut cleanly around the vehicles on it. Stage 2 has flagged the box tumbling off the flatbed in red, along with the smaller fragments already scattered on the asphalt. None of these objects come from a training taxonomy. The pipeline flags them because they sit on the road and are not road.
SAM and SegFormer isolate the drivable road surface from the camera frame, lane by lane, in changing light and weather.
Anything on the segmented surface that isn't road gets flagged, whether a person, a box, a cone, or a tire. The stage is open-set, so the object doesn't need to be in any training taxonomy.
Real debris footage is rare by definition. CARLA on Unreal Engine 5.5, SAM 3D, and Cosmos generate the long tail, labeled scenes of things that almost never happen on camera.
Built on SAM, SegFormer, and Cosmos, and currently in pilot on road-monitoring footage. The same corridor-plus-anomaly pattern transfers to conveyor belts, rail tracks, and runways.
Talk about your corridorThe detectors and the pipeline are not separate tricks. Both come from the house method that already runs our language and physics lines. Adapt a proven base, keep what it knows, and ship the weights with the data.
YOLOX for detection and SAM and SegFormer for segmentation. Published, benchmarked bases, never an architecture from scratch.
Your camera footage, labeled, and extended with synthetic scenes from CARLA UE 5.5 and Cosmos where real examples are too rare to collect.
New classes join the detector while COCO stays in the mix. The cone checkpoints keep the original 80 classes at full strength.
Checkpoint and labeled dataset together, in the size your hardware wants, from a Jetson at the camera to a server GPU.
The cone class isn't an academic demo. It's the same object family running on Simvera, our industrial perception platform, where detectors trained purely in simulation deploy to real factory and road cameras.
Photogrammetry-captured objects rendered into synthetic scenes, with no manual labeling and no waiting months for rare events to happen in front of a camera.
At Simvera's industrial testbed, a detector trained entirely on synthetic data localizes real traffic cones to 6.8px mean accuracy on live camera feeds. The domain gap, closed.
The cone checkpoints prove the recipe. Your defect taxonomy, corridor, or camera network is a labeling campaign away.
Book a technical call
We use cookies
Some cookies are needed for this site to work. We would also like to set analytics cookies to understand how the site is used — but only if you agree. Privacy policy
Strictly necessary
Session and security cookies that keep forms and logins working. These cannot be switched off.
Analytics
Google Analytics, to count visits and see which pages are read. Off unless you turn it on.