omarcevi.dev

ML Inference Systems Engineer · Istanbul, Turkey · open to collaborations

Omar Elcircevi

Building production ML inference systems. Google Developer Expert, Cloud AI · Google Dev Group organizer · Speaker · OSS contributor

inference.sim · one H100, 3 MIG slices · clock slowed 10×
p50 – p99 – throughput – gpu util – queue 0

A queueing model of my day job: requests get dynamically batched onto MIG slices of one H100. Send a burst, then drop max batch to 1 and watch p99. Service times are modeled, not measured.

readabout.md–

About

I work on ML inference systems: deploying and optimizing the platforms that serve recommendation engines, search ranking, and increasingly AI agents in production. Most recently that meant a platform serving 100M+ users at over 1M requests per minute.

My work sits at the intersection of ML, systems, and infrastructure — vLLM and Triton for serving, and Kubernetes across cloud and on-prem. I care about the parts of ML that don't make it into most conference talks: p99 latency, GPU utilization, what breaks at 3am.

I'm a Google Developer Expert in Cloud AI. I co-organize GDG Istanbul and speak regularly at DevFest, Build with AI, and GDG events across Turkey and abroad, in both Turkish and English.

heads 0–11 →avg

These are the real attention weights of GPT-2 small reading the three paragraphs above, pooled from tokens to words. GPT-2 only looks backwards, so arcs always point to earlier words.

Each square is one head; rows are layers 0–11, and brighter squares look further back. Click one to switch.

computed by scripts/ml_figures.py · <|endoftext|> sink removed, rows renormalized

Google Developer Experts 2026 badge — Google Cloud: Cloud AI
readexperience.md–

Experience

readtalks.md–

Talks & Workshops

Selected talks

  • Sep 2026
    Managing Agentic Systems on Google Cloud: From Deployment to Monitoring
    Google Developer Day
  • Apr 2026
    Build with AI: Gemini SDK & Cloud Run
    42 Türkiye
  • Apr 2026
    Build with AI: Google Antigravity
    42 Türkiye
  • Nov 2025
    Architecture of High Performance ML Solutions
    GDG Baku
  • 2023
  • 2023

Workshops

  • Sep 2026
    Architecting Multi-Agent Systems with Google ADK
    Google Developer Day
  • Aug 2026
    Gemma Train-the-Trainer: Running AI Models
    Google Developer Experts
  • May 2026
    Agentverse: Multi-Agent Systems with ADK, MCP and A2A

    A team of research, judge and builder agents talking over A2A, with a self-hosted Gemma model on Cloud Run GPUs.

readprojects.md–

Projects

replay · sdlc-agent-pipeline
t 0.0s cost $0.0000 tool calls 0 tokens 0 → 0 outcome running

A recorded run of my issue-to-PR agents, replayed with its real timestamps, tokens and cost. All recorded runs →

Featured
Agents that turn a GitHub issue into a tested pull request

Planner, coder and reviewer agents in an ADK graph, routed by real diffs and test results instead of model claims. Commands run in sandboxes with no credentials or network, every run is cost-capped, and a human approves before a pull request opens. Deployed on Agent Runtime with keyless CI/CD. Watch recorded runs · architecture

PythonGoogle ADKAgent RuntimeVertex AITerraformGitHub Actions
Load testing vllm-metal on an M1 Pro

Four 4-bit models under AIPerf at 1 to 16 concurrent users. Bursty traffic made time-to-first-token about 5x worse than steady traffic at the same average rate, and the p99 knee came early. Scripts and raw per-request data included.

vllm-metalAIPerfGemmaQwenApple Silicon
Your first model on Triton Inference Server

Serve a ResNet-18 with dynamic batching and two model instances, load test it with perf_analyzer, and deploy it to GKE. Runs on a laptop CPU. Companion code for my intro article.

NVIDIA TritonONNXDockerGKE
A terminal-native AI research digest

Collects new AI papers and posts every day, ranks them by what you like (with or without an LLM), and lets you read full papers without leaving the terminal. Includes hybrid search and Q&A over everything you've read.

PythonGoogle ADKpre-alphaApache-2.0
Turkish meme reactions for Claude Code

A Claude Code plugin (goygoy): hand Claude a task and it answers "Başımla beraber abi! 🫡" — out loud, or in its own reply.

Claude Code pluginhooksbash
readwriting.md–

Writing

All posts on Medium →

readoss.md–

Open Source

readcontact.md–

Elsewhere

I'm open to collaborations: talks and workshops, open-source work, and teams figuring out how to serve models or agents in production. Or just write to compare notes on inference at scale.