City Sample Vehicle Keypoints - 24-point (synthetic-only)

A YOLO-pose model trained entirely on synthetic data - the City Sample 24-point vehicle-keypoint dataset rendered in Unreal Engine 5. It predicts a 24-point anatomical keypoint schema (wheels, head/tail lights, exhaust, roof corners, center, mirrors, bumper and window corners) plus a bounding box per vehicle.

Generated by kiselyovd/ue5-vehicle-synth.

What this model is for

This is a research / proof-of-concept model that demonstrates the synthetic dataset is clean and learnable: a model trained only on it localises vehicles and their keypoints well on held-out synthetic frames. For real-world 14-point vehicle keypoints, see the production model kiselyovd/vehicle-keypoints.

The training labels are derived geometrically from each vehicle's own mesh data (skeletal wheel bones, vertex-cloud roofline, light/mirror material sections), so keypoints are exact per body shape - a van's high tail lights, a pickup's cab-only roof - not a scaled sedan template.

In-domain results (held-out synthetic val)

Metric Box Pose
mAP@50 0.840 0.602
mAP@50-95 0.629 0.390

(Pose mAP is understated: ultralytics uses default OKS sigmas, which are tuned for 17-point human pose, not this 24-point vehicle schema.) Trained from yolo26n-pose on 2,259 synthetic frames (the mesh-derived 2,510-frame capture, 90/10 split), 100 epochs, imgsz 480.

Usage

from huggingface_hub import hf_hub_download
from ultralytics import YOLO

w = hf_hub_download("kiselyovd/citysample-vehicle-keypoints-24pt", "best.pt")
model = YOLO(w)
results = model.predict("your_street_scene.jpg")
# results[0].keypoints.xy -> (N, 24, 2) keypoints per detected vehicle

Honest caveats

  • Synthetic domain. Trained only on rendered frames; the sim-to-real gap is real and measured: on the CarFusion test set this model reaches PCK@0.05 ~0.14 (first 14 points) - it learns vehicle topology but localisation on real photos is unreliable. A label-precision control (re-deriving all keypoints exactly from mesh geometry) left that number unchanged, so the gap is appearance and fleet coverage, not label quality.
  • Fleet coverage. City Sample has on the order of a dozen vehicle models; errors on real photos concentrate on body shapes far from that fleet. See the sim-to-real study for the full analysis.

License

MIT for the weights. Rendered training frames come from Epic's City Sample under the UE EULA (non-interactive renders are distributable; no Epic assets are shipped).

Downloads last month
21
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for kiselyovd/citysample-vehicle-keypoints-24pt

Finetuned
(130)
this model

Dataset used to train kiselyovd/citysample-vehicle-keypoints-24pt