Reliable learning from non-stationary sequential data
How should models learn and generalize when temporal regimes change, observations arrive sequentially, and apparently strong patterns may not survive out-of-time evaluation?
Financial ML · Quantitative Research · Reliable AI
Data Science undergraduate at HKUST(GZ) and quantitative research intern. I study how learning systems remain reliable when data distributions, information sets, and real execution conditions change.
My primary research and career axis is Financial ML / AI-related Quant. I use finance as a demanding setting for broader questions in non-stationary sequential learning, robust evaluation, and reliable decision-making, while keeping reliable autonomous research and data agents as a closely related systems direction.
questions before projects
I organize my work around a small set of methodological questions. Financial markets are the main application domain, but the underlying questions—distribution shift, information boundaries, robust evaluation, and trustworthy evidence aggregation—are broader than finance.
How should models learn and generalize when temporal regimes change, observations arrive sequentially, and apparently strong patterns may not survive out-of-time evaluation?
How can market microstructure, factor/neural features, and predictive models be evaluated under realistic information constraints, regime changes, signal redundancy, and execution assumptions?
How can autonomous systems expose the evidence, state, verification, calibration, and failure semantics required for research workflows to remain auditable?
research & career direction
For graduate study, I want deeper mathematical and methodological training around sequential learning, distribution shift, robust evaluation, and decision-making under changing information. Finance is my primary long-term application and career direction, but I want the research itself to remain general enough to transfer across sequential and data-intensive domains.
Long term, I aim to work on Financial ML / AI-related Quant and the research systems that make high-stakes modeling more reliable.
trajectory in four pieces
Co-author on financial time-series pretraining. My contribution centered on large-scale financial data preprocessing and the research infrastructure required to make terabyte-scale sequence modeling practical.
What it changed for me: moved my interest from isolated forecasting models toward representation learning, scaling, and the information structure of financial sequences.
Independent research with real A-share Level-2 / minute-level data: factor and neural features, temporal/cross-sectional evaluation, signal redundancy, non-stationarity, and walk-forward / out-of-time validation.
What it changed for me: made causal information boundaries, regime robustness, and evaluation design central research concerns rather than post-hoc checks.
Studying evidence aggregation, verification, calibration, and reliability when multiple agents reason over noisy financial information.
Research bridge: connects financial forecasting with a more general question—how autonomous systems should combine imperfect evidence without hiding disagreement or uncertainty.
Independent diffusion-model research on training-free controllable generation, completed through the full problem formulation → method → implementation → experiment → manuscript → peer-review cycle.
What it demonstrated: independent research ownership beyond my primary financial application domain.
peer-reviewed work
research training
Independent factor/neural-feature research using A-share Level-2 and minute-level data, with emphasis on non-stationarity, temporal and cross-sectional evaluation, signal redundancy, and walk-forward / out-of-time validation.
Built supporting research infrastructure for multi-year minute panels, memory-mapped storage, parallel I/O, and model-ready batch construction.
First Prize in 2026 and Second Prize in 2025; practical experience with HPC, numerical workloads, Linux clusters, parallel execution, and performance-oriented engineering.
Undergraduate exchange study in Japan, complementing my Data Science training at HKUST(GZ).
foundation & distinctions
BSc, Data Science and Big Data Technology · 2023–2027 (expected)
CGA 3.807/4.3 · Major CGA 3.837 · Dean’s List ×3.
Selected foundations: calculus, linear algebra, statistics, optimization, algorithms, reinforcement learning, and machine learning.
supporting evidence
I also maintain public systems projects and contribute upstream fixes around state isolation, provider semantics, failure behavior, reproducibility, and release correctness. This is supporting evidence for how I build and validate research systems; the detailed record belongs on GitHub rather than this academic profile.