PORTAL Annual Retreat 2026
Oceano Hotel, Half Moon Bay
Thursday, August 20- Friday, August 21, 2026
Student Posters
Thursday, August 20
| 3:00pm |
Hotel check-in |
|
|
| 4:00-4:15pm |
Welcome - Introduction |
Mark Horowitz |
Faculty |
| 4:15-4:30pm |
Research Overview |
Fred Kjolstad |
Faculty |
| 4:30-5:00pm |
Lightning Talks |
Portal Researchers and Students |
|
| 5:00-6:30pm |
Poster Session & Reception (outside) |
Portal Researchers and Students |
|
| 6:30-8:00pm |
Dinner |
All |
|
Friday, August 21
| 7:45am |
Breakfast |
|
|
| 8:30-10:00am |
Session: Collection Languages and Compilers |
| 8:30-9:00am |
Relational Algebra Compilation |
Fred Kjolstad |
Faculty |
| 9:00-9:30am |
Decoupling Data Layouts from Bounding Volume Hierarchies |
Chris Gyurgyik |
PhD Student |
| 9:30-10:00am |
Partitioning Unstructured Sparse Tensor Algebra for Load-Balanced Parallel Execution |
Atharva Chougule |
MS Student |
| 10:00-10:30am |
Break and guest room check-out |
|
|
| 10:30-11:30am |
Keynote: Computer Architecture in the Age of Agentic AI |
Bill Dally |
Chief Scientist and Senior Vice President of Research, NVIDIA |
| 11:30am-12:00pm |
Student-Industry Speed Interactions |
| 12:00-1:15pm |
Lunch |
All |
|
| 1:15-2:30pm |
Session: Emerging Hardware and Generators
|
| 1:15-1:40pm |
Sphinx: A 7nm 8.192 TOPS Edge-LLM Processor with Outlier-Aware W4A4 Microscaling and Fused 2b KV-Cache Decoding |
Jeffrey Yu |
PhD Student |
| 1:40-2:05pm |
Algorithm–Hardware Co-Design of Formatbook-Based Block-Wise Number Formats for Generative Text and Video |
Wonsuk Jang |
PhD Student |
| 2:05-2:30pm |
Agate: A Design Space Exploration System for Heterogeneous CGRAs Accelerating Dense ML Workloads |
Yuchen Mei |
PhD Student |
| 2:30-3:00pm |
Afternoon Break |
|
|
| 3:00-4:15pm |
Session: ML Compilers and Kernel Languages |
| 3:00-3:25pm |
Scaling Voyager’s ML Compiler to Billion-Parameter LLMs with Automated Tiling, Fusion, and Software Pipelining |
Jeffrey Yu |
PhD Student |
| 3:25-3:50pm |
Extending Voyager for DCIM-Based Accelerator Generation and Compilation |
Allen Pan |
PhD Student |
| 3:50-4:15pm |
Cyclotron: Recurrence Language for Interprocessor Communication |
Shiv Sundram |
PhD Student |
| 4:15-4:45pm |
Explicit Feedback Session |
|
|
| 4:45-5:00pm |
Wrap up and closing thoughts |
Mark Horowitz |
Faculty |
Student Posters
| 1 |
Bobby Yan |
The Right Kernel Every Time: Adaptive Sparse Compilation for Deep Learning |
| 2 |
Rubens Lacouture |
Programming Systems for Sparse Machine Learning on Modern Hardware |
| 3 |
Alexander Root |
Compiling Portable and Performant Spatial Queries with Data and Schedule Independence |
| 4 |
Sai Gautham Ravipati |
Portable Dynamic Tiling for Sparse Tensor Applications |
| 5 |
Atharva Chougule |
Partitioning Unstructured Sparse Tensor Algebra for Load-Balanced Parallel Execution |
| 6 |
Usman Tariq |
Fusion and Tiling for Computation Graphs |
| 7 |
Shiv Sundram |
Cyclotron: Specifying Dataflow in Distributed Architectures |
| 8 |
Chris Gyurgyik |
Data Layout Optimizations via Rewrite Rules |
| 9 |
Anderson Truong |
Flexible analytical performance modeling for deep learning on hierarchical, heterogeneous hardware |
| 10 |
Yasmine Omri |
Agent Memory: System Implications and Co-design for Multi-Tenant Long-Horizon Workloads |
| 11 |
Michael Oduoza |
HW/SW Co-design for SSMs at the Edge |
| 12 |
Allen Pan |
Extending Voyager for Digital Compute-in-Memory-Based Accelerator Generation and Compilation |
| 13 |
Bo Wun Cheng |
Interplay of Activation Quantization and Sparsification for LLM Compression |
| 14 |
Wonsuk Jang |
SemanticDialect: Semantic-Aware Mixed-Format Quantization for Video Diffusion Transformers |
| 15 |
Jeffrey Yu |
Sphinx: An Edge-LLM Accelerator with Outlier-Aware W4A4 Microscaling and Fused 2b KV-Cache Decoding |
| 16 |
Christian Kubicka |
uVLA: Multi-Chiplet 16nm SoP with HMoP for Physical AI |
| 17 |
Yuchen Mei |
Agate: A Design Space Exploration System for Heterogeneous CGRAs Accelerating Dense ML Workloads |
| 18 |
Po-Han Chen |
SPEC: A Scalable and Physical-Design-Friendly Architecture for CGRAs |
| 19 |
Áron Ricardo Perez-Lopez |
Pono 2.0: A Versatile SMT-Based Model Checker for Safety and Liveness |
| 20 |
Zhouhua Xie |
Provably Correct Low-Precision Approximations of Nonlinear Functions for AI Hardware |
| 21 |
Elizaveta Pertseva |
Automated Geometric Predicate Synthesis |
| 22 |
Ritvik Sharma |
CoTenN: Constrained Optimization with Tensor Networks |