More WorkMore Work about Urban Spatial Intelligence

CityRiSE: Reasoning Urban Socio-Economic Status inLarge Vision-Language Models via Reinforcement Learning

Tianhui Liu1,2, Hetian Pang2, Xin Zhang2, Jie Feng2,3†, Pan Hui1†, and Yong Li2†
1 Information Hub, The Hong Kong University of Science and Technology (Guangzhou)
2 Department of Electronic Engineering, BNRist, Tsinghua University
3 Zhongguancun Academy

† Corresponding authors

ACM Multimedia 2026

Urban socio-economic sensing plays a vital role in advancing global sustainable development goals. With the advent of Large Vision-Language Models (LVLMs), new opportunities have emerged to address this challenge by framing it as a multi-modal perception and reasoning task. However, recent studies show that LVLMs still struggle to make accurate and interpretable socio-economic predictions from visual data. To overcome these limitations and fully exploit the potential of LVLMs, we propose CityRiSE, a novel framework for Reasoning urban Socio-Economic status in LVLMs via reinforcement learning (RL). With carefully curated multi-modal dataset and verifiable reward design, our approach guides the LVLM to focus on semantically meaningful visual cues, enabling structured and goal-oriented reasoning for generalist socio-economic status prediction. Experiments demonstrate that CityRiSE, equipped with emergent reasoning, significantly outperforms existing baselines, improving both prediction accuracy and generalization across diverse urban contexts, especially on unseen cities and unseen indicators. This work highlights the promise of combining RL and LVLMs for interpretable and generalist urban socio-economic sensing.

CityRiSE unifies heterogeneous socioeconomic indicators on a common 1–10 scale and trains a generalist urban predictor with GRPO. The framework combines a primary indicator dataset with auxiliary perceptual and general visual reasoning tasks.

Coverage17 global cities

Geographically diverse training and held-out city splits.

Targets10 evaluation tasks

Across in-domain, unseen-city, and unseen-indicator settings.

Efficiency5,109 samples

Strong results without large-scale supervised training data.

OptimizationRL only

GRPO induces reasoning without human-authored chains of thought.

CityRiSE dataset suite and GRPO training framework with keyword and regression rewards
CityRiSE integrates three types of training data with a GRPO pipeline guided by Keyword Reward and Regression Reward.

One model, diverse cities and indicators.

In-domainUnseen CitiesUnseen IndicatorsSatellite ImageryStreet View ImageryInterpretable Reasoning

Three complementary data sources develop task knowledge, urban perception, and transferable visual reasoning. Two verifiable rewards align the model’s reasoning with both semantic evidence and numerical accuracy.

01

Socio-Economic Indicator Data

One satellite image and ten street-view images represent each region. Targets from different indicators are discretized to a unified 1–10 prediction scale.

02

Perceptual Urban Reasoning

Spatial comparison, city geolocation, and socioeconomic ranking teach the model to connect visible urban cues with place and status.

03

General Visual Reasoning

Object counting and pattern completion strengthen foundational perception and abstract reasoning that transfer to urban scenes.

Verifiable, task-aligned rewards

Regression Reward uses a Huber-loss-based signal, rewarding close predictions more than large errors. Keyword Reward encourages task-relevant visual concepts, valid output structure, and location-aware reasoning.

Together they let reasoning emerge through optimization rather than imitation of fixed, human-written templates.

FormatVisual keywordsLocationHuber regression

CityRiSE is evaluated through comprehensive experiments on predictive performance, cross-city and cross-indicator generalization, reward and data design, multimodal inputs, and emergent reasoning.

Tianhui Liu, Hetian Pang, Xin Zhang, Jie Feng, Pan Hui, and Yong Li. CityRiSE: Reasoning Urban Socio-Economic Status in Large Vision-Language Models via Reinforcement Learning. Proceedings of the 34th ACM International Conference on Multimedia (MM ’26), 2026.

@inproceedings{liu2026cityrise,
  title={CityRiSE: Reasoning Urban Socio-Economic Status in Large Vision-Language Models via Reinforcement Learning},
  author={Liu, Tianhui and Pang, Hetian and Zhang, Xin and Feng, Jie and Hui, Pan and Li, Yong},
  booktitle={Proceedings of the 34th ACM International Conference on Multimedia},
  year={2026},
  doi={10.1145/3767308.3835986}
}