City as a separate domain
An object detector trained on camera footage from one city loses accuracy noticeably and predictably when moved to another city. The reason isn't the model itself, but the fact that literally everything changes at once: lighting and white balance of the cameras, traffic density, road markings, the usual car body colors, and even the appearance of typical vehicle models. Formally, this is called geographic domain shift, and in practice it's a recurring headache for those deploying traffic monitoring systems across multiple regions at once.
The classic solution is domain adaptation. But it has an unpleasant property: it almost always requires either fine-tuning of the architecture, which falls apart when hyperparameters change, or direct access to the target city's data — profiling it, tuning thresholds on it, analyzing statistics. And that's where constraints of a decidedly non-technical nature come into play.
What a "blind" training loop means
In ecosystems where data is treated as sensitive information, both familiar paths are closed by definition. You can't export frames from the target city to a research testbed, you can't build statistics from them, you can't tune the model on them. All that remains is a fully "blind" mode: training and evaluation proceed without a single target example, and quality can only be verified where labels already exist and have never left the secure perimeter.
It's precisely within this framework that the authors of a new paper set their task. They're interested not in yet another clever adaptation layer, but in a simpler and more interesting question: how much can be squeezed out of what's already there — out of pretraining and augmentation — if you forgo any manipulation of target data.

Two orthogonal pillars
The proposed training scheme for object detection is built on two independent mechanisms. The word "orthogonal" is key here: they don't duplicate each other and they suppress different sources of error — one works with the semantics and structure of the object, the other with what the model's attention latches onto.
Pretraining on multiple datasets with objectness distillation
The first pillar is a pretraining strategy on several datasets at once with so-called class-agnostic objectness distillation. The idea is to decouple two things that are glued together in ordinary training: the structural geometry of a vehicle and semantic taxonomies.
Put simply, the model first learns to answer the question "is there an object of a characteristic shape here?" without going into how that object is named in a particular dataset. This matters because taxonomies differ between datasets and between cities: somewhere there's a separate class for vans, somewhere everything is lumped into "vehicle," and the boundaries between passenger and cargo vehicles are defined differently. If the model relies on stable geometry rather than local labeling conventions, a change of city hits it less hard.
Grayworld: forcing attention to give up color
The second pillar is a stream of domain-robust augmentation, into which a new transformation called Grayworld has been added. Its job is to pull the rug out from under the model where it has gotten used to cheating.
Detectors often find "shortcuts": they associate an object's class with color. A yellow car is a taxi, a red silhouette is a certain type of vehicle, a specific shade of sky or asphalt is a sign of a particular camera or time of day. Within a single city such cues work, but when transferred to another region they fall apart, because color rendition, lighting, and customary colors there are different. Grayworld suppresses precisely these chromatic labels, forcing the global attention heads to rely on stable priors about shape. Color stops being evidence; geometry remains.

Experiments and results
Evaluation was conducted on the real-time transformer detector RF-DETR. The authors specifically emphasize that the entire setup fits within a modest GPU memory budget — 16 GB, meaning it's reproducible on an ordinary workstation, not just on a cluster.
The optimized variants RF-DETR-HR and RF-DETR-Grayworld delivered a gain of +24.29 over the baseline and took first place on the AI City Challenge Track 6 leaderboard with a score of 47.53 mAP. This is no longer "slightly better than baseline" — it's the difference between a detector that's nearly useless in a new city and one that works there.
The work was accepted to the AI City Challenge workshop at the ECCV 2026 conference; the preprint is available as arXiv:2608.24154 (cs.CV, cs.AI), and the code and data are posted in the SKKUAutoLab/aic26_cross_city repository.
What practitioners should take away from this
A few takeaways that are useful beyond the specific benchmark.
- Augmentation isn't cosmetic. It can deliberately break unwanted correlations if you understand what exactly the model latches onto. Grayworld is an example of a narrow but precise tool against color shortcuts.
- Pretraining can be designed for transfer. Class-agnostic objectness and decoupling geometry from taxonomy is a technique that will come in handy wherever classes are described differently across datasets.
- A "blind" protocol is an honest setup, not a limitation. If a method doesn't peek at target data, its result speaks to real robustness rather than the quality of overfitting.
- Memory budget is part of the result. The 16 GB requirement makes the approach applicable in engineering practice, where top-tier hardware isn't always available.
Caution is also warranted: the conclusions were obtained on a single family of detectors and a single competition track, and Grayworld is, after all, a heuristic rather than a guarantee. Geographic shift remains a problem with no single universal solution, but it's now clear that a significant portion of the progress is achieved before the model ever sees a new city.




