Overview

Projects

My research builds pragmatic, resilient data infrastructure for modern data-driven applications. The guiding principle is that strong data infrastructure should be simple in principle, elegant in design, and efficient in execution, while remaining reliable under production-grade constraints — with four cross-cutting commitments: scalability, robustness, cost-efficiency, and deployability. The work spans four directions.

  • Data Organisation

    Learned indexes for disk-resident, updatable systems, and missing-value imputation for tabular data lakes — making data organisation scalable, robust to incomplete data, and ready for trustworthy analytics.

  • Query Optimisation

    AI-driven cardinality estimation and plan selection for vectors, strings, and complex workloads where classical optimisers fail — reducing estimation errors and unstable execution plans.

  • Agentic Data Processing

    Turning natural-language tasks into explicit, traceable, executable pipelines over heterogeneous data, using structured knowledge, hybrid planning, self-correction, and cost control.

  • Industrial Systems

    Production-grade infrastructure deployed in real settings — including a trajectory data system in Alibaba Cloud serving vehicle, vessel, and aviation workloads.