doublewordai
Popular repositories Loading
-
control-layer
control-layer PublicThe world’s fastest AI model gateway (450x less overhead than LiteLLM). Unified access to LLMs across endpoints (openAI, self-hosted, etc.) behind a single authentication layer - with API key gener…
-
autobatcher
autobatcher PublicDrop-in AsyncOpenAI replacement that transparently batches requests
-
deepseek-reddit-agent
deepseek-reddit-agent PublicAn example notebook which shows how you can build a LLM agent that scrapes information from Reddit and summarize key bullets using a self-hosted DeepSeek-R1-Distill-Llama-8B deployed with Titan Tak…
-
inference-stack
inference-stack PublicThe Doubleword Inference Stack is the easiest & most performant way to run genAI infrastructure in your private environment.
-
inference-lab
inference-lab PublicHigh-performance LLM inference simulator for analyzing serving systems
Repositories
-
- control-layer Public
The world’s fastest AI model gateway (450x less overhead than LiteLLM). Unified access to LLMs across endpoints (openAI, self-hosted, etc.) behind a single authentication layer - with API key generation, user management, request logging, and more
- snapshot Public Forked from ai-dynamo/snapshot
Snapshot gets GPU pods ready in seconds instead of minutes — by restoring a fully initialized GPU worker instead of starting one from scratch. Kubernetes-native checkpoint & restore that runs alongside your existing stack.
-
-
- modelexpress Public Forked from ai-dynamo/modelexpress
Model Express is a Rust-based component meant to be placed next to existing model inference systems to speed up their startup times and improve overall performance.
Most used topics
Loading…