Instructions to use kiselyovd/citysample-vehicle-keypoints-24pt with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- ultralytics
How to use kiselyovd/citysample-vehicle-keypoints-24pt with ultralytics:
from huggingface_hub import hf_hub_download from ultralytics import YOLO # pick the weights file from this repo's "Files and versions" tab weights = hf_hub_download("kiselyovd/citysample-vehicle-keypoints-24pt", "<weights>.pt") model = YOLO(weights) source = 'http://images.cocodataset.org/val2017/000000039769.jpg' model.predict(source=source, save=True) - Notebooks
- Google Colab
- Kaggle
City Sample Vehicle Keypoints - 24-point (synthetic-only)
A YOLO-pose model trained entirely on synthetic data - the City Sample 24-point vehicle-keypoint dataset rendered in Unreal Engine 5. It predicts a 24-point anatomical keypoint schema (wheels, head/tail lights, exhaust, roof corners, center, mirrors, bumper and window corners) plus a bounding box per vehicle.
Generated by kiselyovd/ue5-vehicle-synth.
What this model is for
This is a research / proof-of-concept model that demonstrates the synthetic dataset is clean and learnable: a model trained only on it localises vehicles and their keypoints well on held-out synthetic frames. For real-world 14-point vehicle keypoints, see the production model kiselyovd/vehicle-keypoints.
The training labels are derived geometrically from each vehicle's own mesh data (skeletal wheel bones, vertex-cloud roofline, light/mirror material sections), so keypoints are exact per body shape - a van's high tail lights, a pickup's cab-only roof - not a scaled sedan template.
In-domain results (held-out synthetic val)
| Metric | Box | Pose |
|---|---|---|
| mAP@50 | 0.840 | 0.602 |
| mAP@50-95 | 0.629 | 0.390 |
(Pose mAP is understated: ultralytics uses default OKS sigmas, which are tuned for
17-point human pose, not this 24-point vehicle schema.) Trained from
yolo26n-pose on 2,259 synthetic frames (the mesh-derived 2,510-frame capture,
90/10 split), 100 epochs, imgsz 480.
Usage
from huggingface_hub import hf_hub_download
from ultralytics import YOLO
w = hf_hub_download("kiselyovd/citysample-vehicle-keypoints-24pt", "best.pt")
model = YOLO(w)
results = model.predict("your_street_scene.jpg")
# results[0].keypoints.xy -> (N, 24, 2) keypoints per detected vehicle
Honest caveats
- Synthetic domain. Trained only on rendered frames; the sim-to-real gap is real and measured: on the CarFusion test set this model reaches PCK@0.05 ~0.14 (first 14 points) - it learns vehicle topology but localisation on real photos is unreliable. A label-precision control (re-deriving all keypoints exactly from mesh geometry) left that number unchanged, so the gap is appearance and fleet coverage, not label quality.
- Fleet coverage. City Sample has on the order of a dozen vehicle models; errors on real photos concentrate on body shapes far from that fleet. See the sim-to-real study for the full analysis.
License
MIT for the weights. Rendered training frames come from Epic's City Sample under the UE EULA (non-interactive renders are distributable; no Epic assets are shipped).
- Downloads last month
- 21
Model tree for kiselyovd/citysample-vehicle-keypoints-24pt
Base model
Ultralytics/YOLO26