ASF Member
ASF Member since 2021, with continued participation in Apache community and ALC Beijing activities.
Apache ecosystem
My open-source work centers on data infrastructure that has to survive real production scale.
My Apache work has grown from HBase and Hadoop-era storage infrastructure into Flink state backends, lakehouse formats, remote shuffle, execution acceleration, and AI-era data systems. The common thread is building open infrastructure that is technically deep, community-governed, and proven under demanding production workloads.
ASF Member since 2021, with continued participation in Apache community and ALC Beijing activities.
Apache Flink, HBase, Paimon, Celeborn, Gluten, and Fluss; HBase committer and PMC work dates back to 2016.
Apache Paimon, Apache Celeborn, and Apache Fluss during incubation.
Apache Amoro and Apache GraphAr.
| Layer | Apache Project | Alibaba Product | Status |
|---|---|---|---|
| Streaming Storage | Fluss | Streaming Storage for Apache Fluss* | TLP |
| Lake Format | Paimon | Milvus x DLF Vector Lake | TLP |
| Lakehouse mgmt | Amoro | DLF* | Incubating |
| Shuffle Service | Celeborn | EMR Fusion for Apache Spark | TLP |
| Batch Execution | Gluten | Fusion Kernel | TLP |
| OLAP | Doris* | EMR StarRocks / Stella Kernel | Related ecosystem |
| Streaming | Flink | Realtime Compute for Apache Flink* | TLP |
| NoSQL | HBase | EMR on ECS DataServing | TLP |
| Vector Database | — | Milvus | — |
Items marked with * are adjacent ecosystem references rather than projects I am directly involved in or products I am responsible for.