JD Tech Genius Team(顶尖青年技术天才计划) Researcher
Multimodal foundation-model post-training · Controllable video generation · Streaming video generation.
PhD student at HKUST(GZ)Technical co-founder of LuckyShort
Hi! I’m Xingchen from Sichuan, China—a panda lover interested in general-purpose intelligence: how AI can learn, reason, and act reliably in the real world.
I am a PhD student at CityMind Lab, HKUST (Guangzhou), advised by Dr. Yuxuan Liang. My research connects multimodal agents, reinforcement learning, visual understanding and generation, and urban intelligence.
My research includes first-author ACL and KDD papers and a collaborative NeurIPS Spotlight. Our multi-view generation work was an ACM MM 2026 Grand Challenge Winner (RMB 210,000 prize). I also co-founded ByFusion (LuckyShort), building video agents for commercial production. Research at PixVerse and PCITECH connects my work to video-model post-training and intelligent traffic control.
I see research and entrepreneurship as complementary ways to make AI useful at a broader scale. What motivates me is how intelligence can generalize, adapt, and sustain purposeful work beyond familiar settings.
Connecting perception, reasoning, and action through grounded representations and reinforcement learning.
Long-context video understanding, reference-guided generation, and adaptation that generalizes to new scenes and stories.
Learning from heterogeneous urban data while incorporating spatial structure and physical knowledge.
Winner of the ACM MM 2026 Grand Challenge multi-view generation track — RMB 210,000 prize.
Traffic-R1 was accepted by ACL 2026. Our work connects language-model reasoning with traffic signal control.
I worked with PixVerse on reference-guided video generation, post-training, and evaluation for multi-shot storytelling.
Received the Best Poster Award at the “AI for Society” competition celebrating HKUST's 35th anniversary.
Our collaborative work AgentSense was accepted by WWW 2026.
Our collaborative work was accepted by NeurIPS 2025 as a Spotlight paper. Congratulations to Siru!
I co-founded ByFusion (LuckyShort) and began my PhD at HKUST(GZ), continuing to bring research and video production systems together.
Our collaborative work was accepted by the EMNLP 2025 main conference. Congratulations to Yuhao!
Traffic-R1 received over 1,000 likes and 100,000 impressions on X.
GeoHG was accepted as a full paper at SIGSPATIAL 2025. Thanks to all my collaborators!
DeepUHI won First Prize at the UGOD AI for Urban seminar.
DeepUHI was accepted by KDD 2025 research track. Thanks to my collaborators.
Our collaborative work was accepted by WWW 2025. Congratulations to Xixuan!
I won the First Prize in the HKUST(GZ) thesis writing competition.
Our urban computing survey was accepted by Information Fusion (SCI IF 14.9)
We released our survey on multimodal deep learning for urban computing.
Our team at JBOT Limited began receiving support through HKSTP programmes.
I joined the Hong Kong Center for Construction Robotics (HKCRC) as a research intern.
I joined the Data Acquisition & Analysis Laboratory at HKU as a research intern, supported by HKSAR.
I graduated from Wuhan University with distinction!
Multimodal foundation-model post-training · Controllable video generation · Streaming video generation.
Video agents for long-form storytelling: planning, reference asset coordination, controllable generation, and visual quality evaluation. Related work includes the open-source ViMax collaboration with HKUDS.
Post-training for the V6 / C1 video models, including multi-reference and preference data, reward design, and internal evaluation. A particular focus was identity consistency and coherent transitions between shots.
Led Traffic-R1 and Traffic-VL research on interpretable, generalizable traffic control; contributed to reinforcement learning, multimodal alignment, and post-training for the TransGPT series.
UrbanGPT and its multimodal extension, UrbanGPT-o: joint modeling of temporal signals, language, satellite imagery, and other urban data for spatiotemporal understanding and prediction.
Early multimodal agents for construction: combining language models, computer vision, knowledge retrieval, and grounded engineering data for perception and task execution.
Research in visual generation, reasoning, and the physical world. † Equal contribution. Work under review is marked separately.
ACM MM 2026 · Grand Challenge winner
ACL 2026
Manuscript
KDD 2025 · Oral
ACM SIGSPATIAL 2025 · Oral
Information Fusion
Manuscript · Under review at ACM TOIS
ACM MM Grand Challenge · Winner, multi-view generation · RMB 210,000 prize
Outstanding Team · China Innovation & Entrepreneurship Competition, Jiangsu division
Best Poster Award · HKUST 35th Anniversary “AI for Society”
First Prize · HKUST UGOD AI for Urban seminar
First Prize · HKUST(GZ) academic writing competition
HKSTP Ideation programme support
Outstanding Graduate · Wuhan University
Second Prize · National Structural Design & Information Technology Competition
Excellent Student Scholarships; Outstanding Student and Excellent Student Cadre · Wuhan University
Neurocomputing · Information Fusion
NeurIPS · WWW · ACM MM · AAAI · SIGSPATIAL · KDD · ACL ARR
I enjoy photography, singing, painting, calligraphy, and badminton.
A more personal side ↗Always happy to exchange ideas about agents, visual generation, and AI in the real world.
Get in touch ↗