InferenceX
Open-source platform that continuously benchmarks SGLang, vLLM, and TensorRT-LLM across NVIDIA, AMD, and Google TPU hardware, so results stay current as the stack changes.
Things I've built and worked on.
Open-source platform that continuously benchmarks SGLang, vLLM, and TensorRT-LLM across NVIDIA, AMD, and Google TPU hardware, so results stay current as the stack changes.
Edit Excalidraw diagrams with Claude Code or Codex, review the changes in the same workspace, and export files you can keep editing.
Multi-stage pipeline framework on SGLang for models that take and produce multiple modalities at once, with an OpenAI-compatible server and Real-Time API support.
High-performance serving framework for large language and multimodal models.
Interactive visualization tool for minimizing language deprivation in deaf and hard-of-hearing children.
This site, built with Next.js, TypeScript, and Tailwind CSS. Blog, project showcase, and responsive design.