Skip to content
@doublewordai

doublewordai

Popular repositories Loading

  1. control-layer control-layer Public

    The world’s fastest AI model gateway (450x less overhead than LiteLLM). Unified access to LLMs across endpoints (openAI, self-hosted, etc.) behind a single authentication layer - with API key gener…

    Rust 94 16

  2. autobatcher autobatcher Public

    Drop-in AsyncOpenAI replacement that transparently batches requests

    Python 24 4

  3. deepseek-reddit-agent deepseek-reddit-agent Public

    An example notebook which shows how you can build a LLM agent that scrapes information from Reddit and summarize key bullets using a self-hosted DeepSeek-R1-Distill-Llama-8B deployed with Titan Tak…

    Jupyter Notebook 13 2

  4. inference-stack inference-stack Public

    The Doubleword Inference Stack is the easiest & most performant way to run genAI infrastructure in your private environment.

    Go Template 10 3

  5. zerodp zerodp Public

    ZeroDP implements an efficient zero-copy data parallel approach for serving Mixture-of-Experts (MoE) models, where expert weights are shared across data parallel ranks via CUDA IPC (Inter-Process C…

    Python 10 3

  6. inference-lab inference-lab Public

    High-performance LLM inference simulator for analyzing serving systems

    Rust 9 2

Repositories

Showing 10 of 70 repositories
  • doublewordai/frontend-crates's past year of commit activity
    Rust 0 Apache-2.0 37 0 5 Updated Sep 23, 2026
  • control-layer Public

    The world’s fastest AI model gateway (450x less overhead than LiteLLM). Unified access to LLMs across endpoints (openAI, self-hosted, etc.) behind a single authentication layer - with API key generation, user management, request logging, and more

    doublewordai/control-layer's past year of commit activity
    Rust 94 Apache-2.0 16 50 132 Updated Sep 23, 2026
  • dynamo Public Forked from ai-dynamo/dynamo

    A Datacenter Scale Distributed Inference Serving Framework

    doublewordai/dynamo's past year of commit activity
    Rust 0 1,632 0 26 Updated Sep 23, 2026
  • snapshot Public Forked from ai-dynamo/snapshot

    Snapshot gets GPU pods ready in seconds instead of minutes — by restoring a fully initialized GPU worker instead of starting one from scratch. Kubernetes-native checkpoint & restore that runs alongside your existing stack.

    doublewordai/snapshot's past year of commit activity
    Go 0 Apache-2.0 17 0 5 Updated Sep 23, 2026
  • blog Public
    doublewordai/blog's past year of commit activity
    TypeScript 0 0 0 3 Updated Sep 22, 2026
  • autobatcher Public

    Drop-in AsyncOpenAI replacement that transparently batches requests

    doublewordai/autobatcher's past year of commit activity
    Python 24 MIT 4 2 12 Updated Sep 22, 2026
  • control-layer-chart Public

    A Helm chart for the Doubleword control layer

    doublewordai/control-layer-chart's past year of commit activity
    Go Template 0 Apache-2.0 0 1 11 Updated Sep 22, 2026
  • dw Public
    doublewordai/dw's past year of commit activity
    Rust 4 MIT 0 0 2 Updated Sep 22, 2026
  • uccl Public Forked from uccl-project/uccl

    UCCL is an efficient communication library for GPUs, covering collectives, P2P (e.g., KV cache transfer, RL weight transfer), and EP (e.g., GPU-driven)

    doublewordai/uccl's past year of commit activity
    C++ 0 Apache-2.0 175 0 14 Updated Sep 22, 2026
  • modelexpress Public Forked from ai-dynamo/modelexpress

    Model Express is a Rust-based component meant to be placed next to existing model inference systems to speed up their startup times and improve overall performance.

    doublewordai/modelexpress's past year of commit activity
    Python 0 Apache-2.0 79 0 1 Updated Sep 22, 2026

Most used topics

Loading…