ACM Symposium on Cloud Computing 2026

SoCC'26 will take place in person in Singapore.

November 18-20, 2026

Keynote Speakers

Haibo Chen (Shanghai Jiao Tong University)

Beyond Silicon: Rewiring Operating Systems Abstractions for the Agentic Cloud

Abstract

Cloud computing is undergoing a paradigm shift, transitioning from deterministic microservices to autonomous, dynamic AI agents. These agentic workloads exhibit fundamentally new characteristics: unpredictable execution graphs requiring safe exploration, massive long-lived state, and millisecond-scale responsiveness. Traditional operating system abstractions, such as static containers, block storage, and rigid accelerator scheduling, were not designed for this extreme heterogeneity, leading to severe resource fragmentation and scalability bottlenecks in modern cloud infrastructure. In this talk, I argue that the Agentic Era demands a fundamental rewiring of system-level abstractions. I will present a vision for the Agentic Cloud OS, built on three directions: (1) Agile Compute and Safe Execution: replacing heavy containers with millisecond-level, rollbackable sandboxes; unifying preemptive scheduling across heterogeneous accelerators; and enabling O(1) live autoscaling for large models. (2) Next-Generation Storage and Generative Systems: moving from manually coded file systems to generative, specification-driven architectures; introducing semantic data abstractions; and exploiting ultra-high-density biochemical memory for extreme-scale state persistence of cold storage. (3) AI-Native Orchestration: elevating dynamic agent execution graphs, multi-agent concurrency, and physical/biochemical workflows to first-class cloud resources, enabling the seamless orchestration of autonomous laboratories and cross-domain pipelines. I will conclude by outlining open challenges for the systems community in bridging silicon computing, generative AI, and synthetic biology for the next decade of intelligent infrastructure.

Bio

Haibo Chen is a Distinguished Professor at Shanghai Jiao Tong University, where he founded and directs the Institute of Parallel and Distributed Systems (IPADS). His main research areas include operating systems, distributed systems, machine learning systems, and AI for Science. He has published extensively in top-tier conferences and journals, including Nature, SOSP, OSDI, ISCA, ASPLOS, FAST, EuroSys, and ATC. His work has been recognized with numerous accolades, including Best Paper Awards at SOSP, ASPLOS, EuroSys, and FAST; a Test of Time Award from DSN; a Best Paper Honorable Mention and Research Highlight Award from SIGMOD; a Research Highlight from Communications of the ACM (CACM); and an Honorable Mention for the ACM SIGOPS Dennis M. Ritchie Thesis Award (as Advisor).

He currently serves as an Editorial Board Member for CACM Research and Advances, co-chairs the CACM Regional Special Sections, and is the General Co-Chair for ASPLOS 2028. He is also the founding Chair of the Technical Steering Committee for OpenHarmony, an open-source operating system deployed on hundreds of millions of devices. Prof. Chen is an ACM Fellow and an IEEE Fellow, and currently chairs ACM SIGOPS.



Joseph M. Hellerstein (UC Berkeley and AWS)

From Heisenbugs to Determination: Foundations for Distributed Safety

Abstract

We are entering an era in which the software we depend on is increasingly distributed, just as AI coding agents are making software generation dramatically easier than software assurance. This raises the stakes for making distributed systems trustworthy. Distributed systems bring unique challenges because they are rife with nondeterminism. Messages race, retries duplicate work, failures expose unexpected interleavings, and programs that pass thousands of tests can still harbor Heisenbugs that appear only in production. We currently address these problems with a hodgepodge of seemingly disparate ideas: consistency and isolation levels, coordination protocols, type systems, model checking and fuzz testing. I will argue that these all address a common formal core: determination. Many distributed computations are not simply computing a function; they must choose among multiple admissible outcomes. Determination is the process by which a system irrevocably rules out possible futures in order to make a choice. Determination Theory studies this process directly from a system's specification. At its coordination-free boundary, it yields complete generalizations of two familiar results: CALM, characterizing exactly when coordination is unnecessary, and CAP, characterizing when required coordination becomes incompatible with availability under partitions. Beyond that boundary, the theory measures the irreducible sequential depth of determination work and provides a framework for explaining which commitments determined an observed result. The goal, however, is not theory for its own sake. We are bringing these ideas into Hydro, an open-source Rust framework for distributed systems. Hydro's type system already catches broad classes of nondeterministic distributed behavior at compile time. For properties that cannot be checked at compile time, the semantics of the type system inform Hydro's built-in model checker, allowing it to skip enormous regions of the execution search space whose behaviors are guaranteed equivalent by the types. The same semantics also enable analysis of the consistency properties of observable program outputs. In the talk I will develop these ideas through familiar cloud systems examples and show how a foundational theory of distributed choice can translate into practical tools for building systems that avoid Heisenbugs and use no more coordination than they actually need.