Geographically diverse training and held-out city splits.
CityRiSE: Reasoning Urban Socio-Economic Status inLarge Vision-Language Models via Reinforcement Learning
† Corresponding authors
ACM Multimedia 2026Abstract
Urban socio-economic sensing plays a vital role in advancing global sustainable development goals. With the advent of Large Vision-Language Models (LVLMs), new opportunities have emerged to address this challenge by framing it as a multi-modal perception and reasoning task. However, recent studies show that LVLMs still struggle to make accurate and interpretable socio-economic predictions from visual data. To overcome these limitations and fully exploit the potential of LVLMs, we propose CityRiSE, a novel framework for Reasoning urban Socio-Economic status in LVLMs via reinforcement learning (RL). With carefully curated multi-modal dataset and verifiable reward design, our approach guides the LVLM to focus on semantically meaningful visual cues, enabling structured and goal-oriented reasoning for generalist socio-economic status prediction. Experiments demonstrate that CityRiSE, equipped with emergent reasoning, significantly outperforms existing baselines, improving both prediction accuracy and generalization across diverse urban contexts, especially on unseen cities and unseen indicators. This work highlights the promise of combining RL and LVLMs for interpretable and generalist urban socio-economic sensing.
Framework
CityRiSE unifies heterogeneous socioeconomic indicators on a common 1–10 scale and trains a generalist urban predictor with GRPO. The framework combines a primary indicator dataset with auxiliary perceptual and general visual reasoning tasks.
Across in-domain, unseen-city, and unseen-indicator settings.
Strong results without large-scale supervised training data.
GRPO induces reasoning without human-authored chains of thought.

Generalization
One model, diverse cities and indicators.
Training Design
Three complementary data sources develop task knowledge, urban perception, and transferable visual reasoning. Two verifiable rewards align the model’s reasoning with both semantic evidence and numerical accuracy.
Socio-Economic Indicator Data
One satellite image and ten street-view images represent each region. Targets from different indicators are discretized to a unified 1–10 prediction scale.
Perceptual Urban Reasoning
Spatial comparison, city geolocation, and socioeconomic ranking teach the model to connect visible urban cues with place and status.
General Visual Reasoning
Object counting and pattern completion strengthen foundational perception and abstract reasoning that transfer to urban scenes.
Verifiable, task-aligned rewards
Regression Reward uses a Huber-loss-based signal, rewarding close predictions more than large errors. Keyword Reward encourages task-relevant visual concepts, valid output structure, and location-aware reasoning.
Together they let reasoning emerge through optimization rather than imitation of fixed, human-written templates.
Experiments
CityRiSE is evaluated through comprehensive experiments on predictive performance, cross-city and cross-indicator generalization, reward and data design, multimodal inputs, and emergent reasoning.
Citation
Tianhui Liu, Hetian Pang, Xin Zhang, Jie Feng, Pan Hui, and Yong Li. CityRiSE: Reasoning Urban Socio-Economic Status in Large Vision-Language Models via Reinforcement Learning. Proceedings of the 34th ACM International Conference on Multimedia (MM ’26), 2026.
@inproceedings{liu2026cityrise,
title={CityRiSE: Reasoning Urban Socio-Economic Status in Large Vision-Language Models via Reinforcement Learning},
author={Liu, Tianhui and Pang, Hetian and Zhang, Xin and Feng, Jie and Hui, Pan and Li, Yong},
booktitle={Proceedings of the 34th ACM International Conference on Multimedia},
year={2026},
doi={10.1145/3767308.3835986}
}