πŸ‘‹ A Brief Introduction

I am currently a third-year Ph.D. candidate at the Hong Kong University of Science and Technology (HKUST), supervised by @Prof. Guang Zhang from HKUST-GZ and @Dr. Zhong Li from MSRA.

My research revolves around Data-centric Machine Learning, with a primary focus on LLMs. Specifically, my work is organized around two core dimensions β€” data scheduling (selection and curriculum) and data organization and representation β€” as outlined below:

RQ1Data Scheduling (Selection & Curriculum)
Which data, and in what order, should the model be trained on?
HardPT Β· ACL 2023DoGraph Β· ACL 2026DirEct Β· ICML 2026D3 Β· ICML 2026CUBE Β· ICLR 2027 (Ongoing)GapLens Β· ICLR 2027 (Ongoing)
MicrosoftMicrosoftTencent HunyuanTencent HunyuanTsinghua AIRTsinghua AIR
RQ2Data Organization & Representation
How should we represent and organize complex, heterogeneous data?
LENS Β· ICAIF 2025HGAN-SDEs Β· ICASSP 2026MM-NSDEs Β· AAAI 2026HF Pretraining Β· ProductFinRipple Β· ACL 2025Meituan Nutrition KG Β· Product
JoinQuantJoinQuantTsinghua AIRTsinghua AIRπŸ› HKUST
DATAFoundational Datasets, Benchmarks and Survey
The WoW Β· Corpus Β· ACL 2027 (Ongoing)BizCompass Β· Benchmark Β· ACL 2026From Tokens to Intelligence Β· Survey Β· ACL 2027 (Ongoing)

My work has been published at venues including ICML 2026 (one paper selected as πŸ† Spotlight), ACL 2026, ACL 2023, ACL 2025, ICAIF 2025, and ICASSP 2026 (πŸ† Oral). I also serve as a reviewer for leading conferences such as NeurIPS, ICLR, and ICML (selected as πŸ† Gold Reviewer).

πŸ“ Selected Publications

  • [ICML 2026] Yuanjian Xu, et al. D3: Dynamic Directional Graph-Constrained Data Scheduling for LLM Training [ Paper] [ Code]. [CCF A] [CORE A*] (@ MSRA)
    βœ“ Key Contribution: We explain why training order matters in LLM optimization and propose a data scheduling framework grounded in gradient interactions, where training dependencies are modeled as a graph that explicitly constrains valid training orders.
  • [ICML 2026 Spotlight πŸ†] Yuanjian Xu, et al. Towards Efficient LLMs Annealing with Principled Sample Selection [ Paper] [ Code]. [CCF A] [CORE A*] (@ MSRA)
    βœ“ Key Contribution: We provide a theoretical characterization of steady-state properties in LLM annealing and formulate sample selection as an optimization problem, achieving SOTA results across multiple model scales.
  • [ACL 2026] Yuanjian Xu, et al. Rethinking Data Mixing from the Perspective of Large Language Model [ Paper]. [CCF A] [CORE A*]
    βœ“ Key Contribution: We establish formal connections between gradient dynamics and domain distributions, and introduce DoGraph, a graph-constrained optimization framework for data mixing that clarifies how domain weighting influences LLM generalization.
  • [ACL 2023] Yuanjian Xu, et al. Hard Sample Aware Prompt-Tuning [ Paper]. [CCF A] [CORE A*] (@ THU AIR)
    βœ“ Key Contribution: We introduce a hard sample aware mechanism for prompt-tuning that dynamically adjusts learning focus on difficult samples, improving model performance on challenging instances.

Other Publications

  • [ACL 2026] Yuanjian Xu, et al. BizCompass: Benchmarking the Reasoning Capabilities of LLMs in Business Knowledge and Applications [ Paper] [ Code]. [CCF A] [CORE A*]

  • [ACL 2025] Yuanjian Xu, et al. FinRipple: Aligning Large Language Models with Financial Market for Event Ripple Effect Awareness [ Paper] [ Code]. [CCF A] [CORE A*]

  • [ICASSP 2026 Oral πŸ†] Yuanjian Xu, et al. HGAN-SDEs: Learning Neural Stochastic Differential Equations with Hermite-Guided Adversarial Training [ Paper]. [CCF B] [CORE A]

  • [ICAIF 2025] Yuanjian Xu, et al. LENS: Large Pre-trained Transformer for Exploring Financial Time Series Regularities [ Paper]. (Leading conference for AI in Finance)

Under Review

  • [Technical Report @ Tencent Hunyuan] Jinyi Han, Yuanjian Xu, et al. Skill-Use: Can LLMs Actually Use Skills in Agentic Harnesses? [ Paper] [ Code]. [CCF A] (under review at AAAI 2027, equal contribution)

  • [Under Review at AAAI 2027] Yuanjian Xu, et al. Rethinking Neural SDEs under Shifting Data-Generating Processes. [CCF A]

  • [Under Review at NeurIPS 2026] Yuxuan Sun, Yuanjian Xu, et al. Rethinking Knowledge Distillation for Diffusion Language Models. [CCF A] [CORE A*] All Positive Reviews (@ MSRA)

In Progress

  • [🚧 Target ICLR 2027] Yuanjian Xu, et al. Benchmarking Sample Reasoning Quality from Internal Geometric Dynamics in LLMs. [CCF A] (@ Tencent Hunyuan)

  • [🚧 Target ACL 2027] Yuanjian Xu, et al. The WoW: A Large-Scale World Knowledge Corpus for Full-Lifecycle LLM Training. [CCF A]

  • [🚧 Target ACL 2027 Β· Survey] Yuanjian Xu, et al. From Tokens to Intelligence: A Survey on Data Selection for Large Language Models. [CCF A]

πŸŽ“ Education

I am currently pursuing a Ph.D. in Fintech at the Hong Kong University of Science and Technology. I received my Master’s degree in Computer Science from Peking University, and my Bachelor’s degree in Computer Science from Nankai University.

πŸ”¬ Academic Activities

  • Teaching Assistant, Advanced Statistics (FTEC 5030), HKUST

πŸ’Ό Internship

Tencent
Tencent Hunyuan Top Talent Research Intern

Working on quality assessment for unlabeled samples and the design of frontier agentic benchmarks.

Tsinghua University
AIR, Tsinghua University Research Intern

Worked under the supervision of Prof. Zaiqing Nie. Contributed to the Meituan Nutrition Knowledge Graph construction. Investigated hard sample problems in NLP and proposed HardPT, published at ACL 2023.

Microsoft Research Asia
Microsoft Research Asia (MSRA) Research Intern

Worked under the supervision of Dr. Zhong Li. Focused on data selection and training order optimization for large language models. Proposed the D3 method and an annealing training framework, both accepted at ICML 2026 (annealing work as Spotlight).

JoinQuant
Joinquant (Billion-scale quantitative fund) Research Intern

Worked under the supervision of PM Ruixiao. Developed tick-level generative and representation models for high-frequency trading. Addressed key challenges including non-equally spaced data and market randomness.

πŸ† Honors and Awards

  • 2026 Top 10% Intern, Microsoft Research Asia (MSRA)
  • 2023–Present Full Ph.D. Scholarship, Hong Kong University of Science and Technology
  • 2021 Award for Excellent Academic Excellence, Peking University (Certificate No.: H2021000170320)
  • 2021 Air Star Plan, Tsinghua University, Institute for AI Industry Research (AIR)