Professional Summary
Cloud Platform & Site Reliability Engineer with 5+ years of experience in multi-cloud environments (GCP, AWS, Alibaba Cloud, Azure and ByteCloud) and core distributed systems middleware. Proven track record of scaling high-throughput systems and driving engineering efficiency. Proficient in Go, Python, Kubernetes, and Terraform.
Engineering Experience
- Serve as the primary maintainer for 14 core middleware platforms (including storage, caching, and scheduling), delivering 99.99% reliability by leading 24/7 on-call operations and major cross-functional incident responses for Global E-commerce teams distributed across the US, Asia, Europe, and the Rest of the World.
- Built an AI-powered internal assistant using custom language models (Openclaw) to automate root-cause analysis for common system incidents, reducing engineering workload by 50-70% and cutting total triage time by nearly 70%.
- Enhanced operational stability and monitoring for large-scale job scheduling systems, eliminating approximately 100,000 monthly failures and reducing incident recovery times (MTTR) by 25%.
- Drove company-wide engineering initiatives by standardizing service agreements (SLAs), building real-time stability dashboards, and streamlining incident management to cut QA triage time by 40%.
- Orchestrated the zero-downtime migration of 60+ Tokopedia services to ByteDance's internal cloud by building a custom orchestration platform across GCP, AWS, and Alibaba Cloud.
🤖 Architecture Spotlight: AI-Powered On-Call Assistant
- Automated First Responder: Engineered a custom AI bot utilizing Lark Bot and Openclaw to autonomously manage and monitor 14 core middleware platforms. Programmed the bot to seamlessly intervene and initiate diagnostics if an on-call engineer does not respond within 5 minutes of a mention.
- Dynamic Infrastructure Querying: Developed advanced capabilities allowing the AI to securely authenticate via ByteCloud service accounts to dynamically query service monitoring systems, extract infrastructure metrics, and parse system logs in real-time.
- Accelerated Triage: Leveraged custom language models to autonomously analyze incident data and perform root-cause analysis, successfully reducing incident triage time by nearly 70% and significantly cutting overall engineering workload.
- Led the enterprise-wide migration from Jenkins to GitHub Actions, standardizing CI/CD pipelines (including Ansible, Packer, and Terraform) across Tokopedia and its subsidiaries.
- Enabled engineers to develop CI/CD use cases using self-hosted GitHub Runners, significantly enhancing pipeline scalability and maintainability.
- Automated the identification and decommissioning of over-provisioned infrastructure, driving significant cost savings and improving overall resource utilization.
- Handled daily infrastructure operations and performance tuning across GCP, AWS, and Alibaba Cloud. Collaborated with the GoTo-Financial team to successfully implement and migrate Kubernetes for GoTo-Logistic companies.
🚀 Architecture Spotlight: Multi-Cloud CI/CD Automation
- Scalable CI/CD Architecture: Architected enterprise-wide GitHub Actions CI/CD pipelines utilizing self-hosted Kubernetes runners and custom Helm charts to enable dynamic Horizontal Pod Autoscaling (HPA).
- Standardized Tooling: Empowered global engineering teams by deploying organizational-level custom runner images pre-configured with Ansible, Docker, Packer, and Consul, seamlessly connecting staging and beta environments for performance testing.
- GitOps & Deployment: Implemented automated GitOps workflows utilizing self-testing loops and pull request (PR) comment-based deployment approvals.
- Multi-Cloud Image Provisioning: Streamlined infrastructure deployment by automating image builds (via Packer and Ansible) to securely deploy pre-configured application components across GCP, AWS, Alibaba Cloud, and Azure for compute instance provisioning (e.g., EC2, GCE).
- Created and maintained service applications, designing robust product APIs between services for frontend and UI/UX teams according to client requirements. Designed and executed queries for flow processing jobs using KSQL and Kafka.
- Deployed and monitored applications automatically using Jenkins, Kubernetes, and other deployment tools. Implemented comprehensive alert and notification systems using Grafana, Prometheus, and scheduling.
Life at Work
A glimpse into my time collaborating with incredible engineering teams at Tokopedia and Bytedance.
Skills & Competencies
Cloud Platforms
Google Cloud Platform
AWS
Alibaba Cloud
Azure
ByteCloud (Volcano Engine)
Infrastructure & DevOps
Ansible
Terraform
Packer
Kubernetes/Docker
Consul
NGINX
Akamai
CI/CD
Github Actions
Gitlab CI
Jenkins
Programming & Scripting
Go
Python
Java
NodeJS
Data & Middleware
ClickHouse
DataLeap
Kafka
KSQL
Prometheus
Grafana
AI & Dev Productivity
Openclaw
TRAE
Cursor
Languages
English
Indonesia
Education & Certifications
- AWS - Cloud Native, Amazon Web Services (AWS), 10/01/20
- AWS - Building Serverless Applications, AWS, 08/01/20
- Google Cloud Platform Big Data and Machine, Google, 10/01/20
- API Design & Fundamental of Google Cloud's Apigee API Platform, Google, 09/01/20
- AWS Gameday 2023 (Third Place)