Tech Week LogoLos Angeles, October 12 - October 18
LA calendar

LA Tech Week 2026Featured

Self-Host Frontier AI: Running Open-Weight Models on AWS

Hosted by Amazon Web Services

When
1:00 PM
Where
Santa Monica
Format
Roundtable / Workshop
Self-Host Frontier AI:  Running Open-Weight Models on AWS artwork

/ About this event

The open-weight model revolution is here — frontier models matching or beating closed APIs on benchmarks, and you can run them on your own infrastructure. But hosting them in production isn't as simple as pulling weights from HuggingFace.

How much GPU memory do you actually need? What happens to throughput as context windows grow? When does managed infrastructure save you money vs. burn it? And how do you go from "it runs on my machine" to serving real production traffic?

Join us for a 90-minute deep dive where we'll cut through the hype and show you exactly how to deploy the latest open-weight frontier models — from choosing the right hardware to production-grade serving at scale.

What You'll Learn

🏗️ The GPU Math That Matters — Understand the memory tradeoffs that determine whether your deployment handles 1 user or 100 concurrent users. Context length vs. throughput — and how to think about it

🗺️ A Clear Decision Framework — From fully managed endpoints to self-hosted clusters — which path fits your model, your traffic pattern, and your team's ops appetite

💰 The Cost Math Your CFO Will Ask About — Managed APIs vs. self-hosted: when does convenience justify the premium? We'll walk through real break-even scenarios

🔧 Day-0 Model Support — How to run the newest models the day they drop, without waiting for official platform support

🌐 Data Sovereignty & Control — Why some teams are moving off API providers and back to self-hosted — and when it actually makes sense

Models We'll Cover

We'll walk through real deployment scenarios using today's top open-weight models — from lightweight models that fit on a single GPU to the largest frontier models pushing hardware limits:

GLM • Kimi K3 • Gemma • DeepSeek • MiniMax M3 (and whatever drops between now and October)

Covering the full spectrum — from models any team can deploy affordably to trillion-parameter beasts that require serious infrastructure planning.

Who Should Attend

• ML/AI Engineers — Building inference pipelines and choosing between serving frameworks

• Platform Teams — Evaluating managed vs. self-managed GPU infrastructure

• CTOs & Founders — Deciding whether to self-host or use API providers

• Anyone who's Googled "how much VRAM do I need for \[model\]" in the last month

Ideal for teams actively evaluating self-hosted inference — whether for data sovereignty, cost control, customization, or all three.

What You'll Walk Away With

✅ A decision framework for choosing the right deployment path for any open model ✅ Hardware selection guidance — how to right-size your GPU setup ✅ Key considerations for production-grade model serving ✅ Cost comparison: managed APIs vs. self-hosted vs. GPU cloud providers ✅ Direct access to AWS Solutions Architects who deploy these models with startups weekly

Event Details

📅 Date: October 17, 2026 11:00 AM PT (LA Tech Week) ⏱️ Duration: 90 minutes 📍 Location: 2450 Colorado Ave, Santa Monica, CA 90404 👥 Capacity: 80–100 (first come, first served) 💬 Bring: Your toughest "how do I host X model?" questions

Hosted By

AWS Startups — Built by Solutions Architects who deploy open-weight models with startups every week. This isn't slides about what's theoretically possible — it's what we ship in production.

This event is a part of #LATechWeek—a week of events hosted by VCs and startups to bring together the tech ecosystem. Learn more at [www.tech-week.com](http://www.tech-week.com).

/ Stay in the loop

Get the LA Tech Week schedule in your inbox

New events, featured picks and reminders, straight to your inbox.

Tech Week 2026

Questions? Email us at hello@tech-week.com

Download Media/Press KitJoin the a16z speedrun talent networkView Community Guidelines

Follow Our Socials