KCD San Francisco Bay Area 2026

KCD SF Bay Area

Sep 1, 4:00 PM – Sep 2, 2:00 AM (UTC)

Computer History Museum, Mountain View, CA

In-person event

Get tickets

About this event

KCD San Francisco Bay Area is back for 2026!

Tuesday, September 1, 2026 at the Computer History Museum in Mountain View, CA

Join us for the second KCD San Francisco Bay Area, the biggest gathering for the cloud native and Kubernetes community in the Bay Area! KCD San Francisco Bay Area is a full-day event with sessions that explores the latest cloud native technologies and topics like Cloud Native AI, Platform Engineering, Security, Observability, trending cloud native open source projects, and more.


Join us for amazing opportunities for community networking with some of our industry's most iconic thought leaders and technologists.

The event will take place at the historic Computer History Museum in Mountain View, and as an added bonus, the museum will be open during the day to all event attendees as part of your registration package.

https://computerhistory.org/wp-content/uploads/2024/10/2024_10_ChatbotDecoded_Entrance-CS0001-scaled-540x406-c-default.jpg

Cap off the day with our vibrant evening social hour, perfect for connecting with industry friends, colleagues, and new acquaintances, or stroll through the Computer History Museum exhibits during the social hour. Whether you are here to learn, share ideas, or make new connections, this event is designed to foster collaboration and knowledge-sharing within the Bay Area's thriving cloud native community.

Speakers

  • Tamao Nakahara

    Guild.ai

    Director of Growth and Community

  • Paul Zimmerman

    Uber

    Developer Advocate

  • Duffie Cooley

    Isovalent

    Field CTO

  • Julia Furst Morgado

    Dash0

    Principal Developer Relations Engineer

  • Leigh Capili

    ControlPlane

    Principal Cloud Native Consultant

  • Reese Lee

    New Relic

    Senior Developer Relations Engineer

  • Venkat Gattupalli

    OpenAI

    Scaling Infrastructure for Frontier Models

  • Ole Lensmar

    Testkube

    CTO

  • Archana Anand

    Linkedin

    Archana Anand Senior Software Engineer, Kubernetes Infra Team

  • Cijo Thomas

    Microsoft

    Principal Software Engineer

  • Eric Wang

    Uber

    Senior Staff Engineer, ML Platform

  • Archana Kataria

    Intuit

    Group Development Manager

  • Ishan Shah

    PayPal

    Software Engineer | Distributed Systems, AI, and Platform Engineering

  • Garvit Kataria

    Intuit

    Senior Software Engineer

  • Haibing Zhou

    OpenAI

    Engineer

  • AmyJune Hineline

    Linux Foundation

    Certification Community Architect

  • Goutham Annem

    AWS

    Sr Technical Account Manager

  • Priya Namasivayam

    EarnIn

    Staff Platform Engineer

  • Akshay Pratinav

    Intuit

    Senior Staff Software Engineer

  • Sameera Jayasoma

    WSO2

    VP & Distinguished Engineer

  • Ronald Petty

    RX-M

    Principal Consultant

  • Giorgi Keratishvili

    EPAM Systems

    Lead DevOps

When

When

September 1 – 2, 2026
4:00 PM – 2:00 AM (UTC)

Organizers

  • Lisa-Marie Namphy

    Program Chair | CNCF Ambassador | DevRel Architect

  • Rey Lejano

    Red Hat

    Project Manager | SF Chapter Organizer

  • Jason Smith

    Google

    Special Programs | SF Chapter Organizer

  • Matthew Cascio

    American Red Cross

    Sponsors + Finance | CNCF Ambassador

  • Natalie Lunbeck

    Man Numeric

    Program Committee

Platinum

Plural logo

Plural

Akuity logo

Akuity

Gold

Teleport logo

Teleport

Community Partner

KubeEvents logo

KubeEvents

Kubecrash logo

Kubecrash

Merge Forward logo

Merge Forward

Workshop

Guild.ai logo

Guild.ai

Schedule

Welcome Intros & Keynote

The Next Decade of Agents: How Agentic Computing is Reshaping Cloud Native - Ronald Petty, RX-M

Over the past decade, cloud native transformed how we build, deploy, and operate software. Containers, Kubernetes, and the CNCF ecosystem became the foundation of modern applications. Now another shift is underway. AI agents are evolving from chat interfaces into systems that can reason, plan, use tools, collaborate, and take action. As these capabilities mature, an important question emerges: How will agentic computing reshape cloud native over the next decade? This talk briefly looks back at the evolution of cloud native before exploring how agents may influence Kubernetes and the broader CNCF ecosystem. We'll examine trends in agent communication, identity, memory, observability, governance, and orchestration, and discuss which cloud native concepts may endure, evolve, or give way to new abstractions. Rather than predicting specific products, this session focuses on the long-term patterns that could define the next generation of distributed systems.

Code to Cluster: Abstracting Kubernetes ML Complexity with Michelangelo - Paul Zimmerman & Eric Wang, Uber

Building machine learning platforms on Kubernetes shouldn't force data scientists to become infrastructure experts. True platform engineering means masking the complexities of cluster manifests, distributed scaling, and fault tolerance so data teams can focus purely on innovation. In this session, Paul Zimmerman and Eric Wang from Uber's Michelangelo team explore how platform engineers can leverage the newly open-sourced Michelangelo framework to build a seamless "Code to Cluster" experience. We’ll dive into strategies for abstracting infrastructure while retaining native control over Kubernetes primitives, deploying complex stacks via Helm, scaling distributed training with Ray and PyTorch, and using CNCF Cadence Workflow for pipeline resilience. Join us to see how to bridge the gap between abstract code and distributed cloud-native scale.

Agentic GitOps: Agent and Sandbox Guardrails for CI/CD - Tamao Nakahara, Guild.ai & Leigh Capili, ControlPlane

With agentic workflows, kubectl commands can have dangerous consequences. Improper RBAC can allow agents to kubectl delete, and that includes deleting your whole CI with no commits to roll back to! That’s why Flux's security-first design is even more relevant for agentic GitOps. We'll cover how to use Flux to confine agents to human-reviewable PRs for all sorts of use cases. We’ll do this with a kernel-sandboxing tool called `nono`. In addition, for your agent management tool of choice, we'll cover how to manage your agents' sandboxes with Flux so that nefarious (or confused) agents can't destabilize the security policies that you have in place. We’ll cover safe practices for agentic use cases like: - using Prometheus metrics to trigger resource tuning - troubleshooting and rolling back after HPA crashes - agents requesting additional network access with PR’s for human reviewers Come join in!

How did that happen? And is it a security Problem (Lightning Talk & Birds of a Feather) - Duffie Cooley, Isovalent at Cisco

Join Duffie Cooley as we dig in to learn about a curious engineer uncovered one of the most complex social engineering attacks in recent memory. Learn about how it was discovered and how we can use tools to better understand what’s happening on our own systems.

Behind the Exam Curtain: The Secret Life of a Subject Matter Expert (Lightning Talk & Birds of a Feather) - AmyJune Hineline, Linux Foundation

Ever wonder who writes the certification exam questions that keep you on your toes? In this quick and insightful session, we’ll take a peek behind the curtain at how exams are built and the crucial role Subject Matter Experts (SMEs) play in shaping them. Learn what it means to be an SME, how the process works from blueprint to beta, and why volunteering your expertise is both rewarding and career-boosting. Whether you love sharing knowledge or just want to give back to your community, you’ll discover how to help build the future of open-source certifications.

Rocket Power Your Kubernetes Career With Kubestronaut Program (Lightning Talk & Birds of a Feather) - Giorgi Keratishvili, EPAM

Are you a person who wants to fly high? Conquer mountains of Kubernetes certifications then this talk is for you, Giorgi will share all details of kubestronaut program, what benefits does it gives to person and his certification journey as he holds all 5 and even more certificates from CNCF also he has been beta tester and exam developer some of them...

Lunch & Birds of a Feather

OpenChoreo: Developer Platform for both Humans and Agents (Lightning Talk & Birds of a Feather) - Sameera Jayasoma, WSO2

Kubernetes gives platform teams powerful building blocks. But turning those into a real developer experience takes months of work. OpenChoreo is a complete, open-source developer platform for Kubernetes. It's ready to use from day one, for both humans and agents. In this lightning talk, I'll walk through how OpenChoreo provides development and platform abstractions on top of Kubernetes. It comes with a Backstage-powered developer portal, plus built-in CI/CD, GitOps, and observability. Developers can self-serve deployments without needing to be Kubernetes experts. These same abstractions also work well for AI agents. Agents need to build, deploy, and operate workloads too, often alongside humans. I'll share what it means to design a platform that serves both audiences from the start. If you're building a platform for your team, or thinking about how agents fit into your Kubernetes setup, this talk is for you. You'll leave with a clear picture of what a CNCF Sandbox developer platform looks like when built for both humans and agents.

When Packets Disappear in the Cloud: Debugging Kubernetes Inference Workloads - Venkat Gattupalli & Haibing Zhou, OpenAI

Large scale inference workloads on Kubernetes are extremely sensitive to rare packet loss: one missing packet can cause seconds of tail latency or failed requests. These incidents are hardest when they cross the visibility boundary between platform teams and cloud-provider infrastructure: node kernel, vNIC driver, hypervisor, and network. In this session, we share an incident response story from OpenAI's scale inference. Starting from latency symptoms and TCP retransmissions, we used targeted eBPF tracing to follow packets through the node stack and vNIC driver. Correlating sequence numbers, skb state, DMA mappings, and completions showed affected packets leaving the guest driver path, narrowing the issue toward cloud infrastructure. That evidence enabled a shadow traffic workaround, validated latency improvement, and built rollout confidence while the provider fix was underway.

Why Is Everyone Still Sending Raw Data? - Julia Furst Morgado, Dash0 & Reese Lee, New Relic

Most teams running the OTel Collector are using it as a passthrough. Data goes in, data goes out and everything lands in the backend whether it is useful or not. That gets expensive fast and it creates compliance problems when PII ends up in your traces. OTTL, the OpenTelemetry Transformation Language, ships with every Collector distribution and lets you filter, redact, and normalize telemetry data before it ever leaves your infrastructure. In this session we'll look at the syntax and architecture, walk through real transformation statements for the transform and filter processors, and cover the cases where OTTL makes more sense than building a custom component.

Self-Healing Systems: How LLM Agents Are Reinventing Cloud-Native Disaster Recovery - Akshay Pratinav, Garvit Kataria, & Archana Kataria, Intuit

Cloud-native systems have outgrown the incident response playbook. Static runbooks, manual operator intervention, and reactive on-call rotations weren't designed for the scale and complexity of modern distributed systems — and the result is longer outages, inconsistent recoveries, and burned-out engineers. This session introduces an agentic AI approach to disaster recovery, where LLM-based agents don't just advise — they act. Multiple collaborating agents work in concert to detect anomalies, reason about root causes, evaluate recovery strategies, and execute remediation directly through your existing Kubernetes and cloud interfaces. Human oversight remains intact through safety controls for high-impact operations, so you keep the guardrails without the bottlenecks. We'll show real-world results: dramatic reductions in MTTR and operational toil, with measurable improvements in recovery consistency and reliability. More importantly, we'll walk through the architecture — how agents are orchestrated, how they reason under uncertainty, and how this system evolves from a passive advisory tool into an autonomous SRE co-pilot. If you're an SRE, platform engineer, or architect wondering where AI fits into your reliability story, this session gives you a concrete, battle-tested answer.

Shift-Left for Platform Teams: Kubernetes-Native Infrastructure Testing at EarnIn - Priya Namasivayam, EarnIn & Ole Lensmar, Testkube

EarnIn operates a high-availability fintech platform on 15+ CNCF projects, including Flux, Karpenter, Kyverno, cert-manager, Linkerd, external-secrets-operator, Argo CD, Velero, and more. When infrastructure components fail silently — cert-manager chains breaking post-upgrade, Velero jobs reporting success without executing — the impact is immediate and customer-facing. To address this, we built a Kubernetes-native testing framework using Testkube that validates our full infrastructure stack on every GitOps-triggered change. Tests are defined as Kubernetes CRDs, integrated into our Flux delivery pipeline, and cover TLS validation, policy enforcement, secret sync, backup integrity, and add-on health across the stack. This session walks through our test architecture, the specific failure classes we catch, and how we operationalized infrastructure testing as a first-class part of our platform engineering workflow — not an afterthought.

Raffle & Closing Comments

Happy Hour

CONTACT US