Overview
Projects
My research builds pragmatic, resilient data infrastructure for modern data-driven applications. The guiding principle is that strong data infrastructure should be simple in principle, elegant in design, and efficient in execution, while remaining reliable under production-grade constraints — with four cross-cutting commitments: scalability, robustness, cost-efficiency, and deployability. The work spans four directions.
-
Data Organisation →
Learned indexes for disk-resident, updatable systems, and missing-value imputation for tabular data lakes — making data organisation scalable, robust to incomplete data, and ready for trustworthy analytics.
-
Query Optimisation →
AI-driven cardinality estimation and plan selection for vectors, strings, and complex workloads where classical optimisers fail — reducing estimation errors and unstable execution plans.
-
Agentic Data Processing →
Turning natural-language tasks into explicit, traceable, executable pipelines over heterogeneous data, using structured knowledge, hybrid planning, self-correction, and cost control.
-
Industrial Systems →
Production-grade infrastructure deployed in real settings — including a trajectory data system in Alibaba Cloud serving vehicle, vessel, and aviation workloads.