Apache ecosystem

Open infrastructure at production scale.

My open-source work centers on data infrastructure that has to survive real production scale.

Long-Running Threads

My Apache work has grown from HBase and Hadoop-era storage infrastructure into Flink state backends, lakehouse formats, remote shuffle, execution acceleration, and AI-era data systems. The common thread is building open infrastructure that is technically deep, community-governed, and proven under demanding production workloads.

Apache Roles

ASF Member

ASF Member since 2021, with continued participation in Apache community and ALC Beijing activities.

PMC Member

Apache Flink, HBase, Paimon, Celeborn, Gluten, and Fluss; HBase committer and PMC work dates back to 2016.

Champion

Apache Paimon, Apache Celeborn, and Apache Fluss during incubation.

Mentor

Apache Amoro and Apache GraphAr.

Apache Timeline

Open Source Ecosystem

LayerApache ProjectAlibaba ProductStatus
Streaming StorageFlussStreaming Storage for Apache Fluss*TLP
Lake FormatPaimonMilvus x DLF Vector LakeTLP
Lakehouse mgmtAmoroDLF*Incubating
Shuffle ServiceCelebornEMR Fusion for Apache SparkTLP
Batch ExecutionGlutenFusion KernelTLP
OLAPDoris*EMR StarRocks / Stella KernelRelated ecosystem
StreamingFlinkRealtime Compute for Apache Flink*TLP
NoSQLHBaseEMR on ECS DataServingTLP
Vector DatabaseMilvus

Items marked with * are adjacent ecosystem references rather than projects I am directly involved in or products I am responsible for.