Alibaba Cloud EMR

Cloud-scale data products.

Product Lines

Spark / Fusion

Performance-oriented Spark platform work, including Fusion-related acceleration and serverless lakehouse analytics.

StarRocks / Stella

OLAP and lakehouse analytics work around EMR StarRocks and Stella Kernel performance leadership.

Milvus

Vector database work for AI-era retrieval, multimodal data, and vector lake scenarios.

EMR Agent

AI-assisted EMR operations and product experience for modern cloud data platforms.

EMR Scope

DataOps Platform

Product and engineering work for cloud data platform development, operations, and user workflows.

Computing Engines

Flink, Hive, Spark, StarRocks, Trino, and related open-source engines in Alibaba Cloud EMR.

Metadata & Security

Hive Metastore, Ranger, and the management layer needed by enterprise data platforms.

Lake Storage

Cache, lake format, lake file system, and data lake storage engines for lakehouse and AI-era workloads.

EMR Evolution

My EMR work has focused on evolving the platform from open-source component management toward cloud-native and serverless data infrastructure, including EMR 2.0, EMR Serverless Spark, lakehouse analytics, and AI-assisted data platform capabilities.

More recently, this direction has expanded into OpenLake, lake-stream integration, multimodal data storage and access, and the shift from human-centric data access patterns to AI-agent-driven data generation and retrieval.

Earlier Systems Work

Earlier systems work in HBase, search data infrastructure, realtime compute storage, and Flink state backends shaped the current product direction: production-first, open-source-compatible, and measured under real traffic.

Engineering Focus