Sign inSign up

kienxuandaoit/my_aigateway

By kienxuandaoit

•Updated 3 days ago

HTTP AI gateway

Image
Machine learning & AI
0

1.3K

kienxuandaoit/my_aigateway repository overview

⁠my_aigateway (Docker deploy)

Pass-through AI gateway: one Bearer key fans out to OpenAI / Responses / Gemini / Claude / embeddings upstreams. No request translation, no keys baked into the image.

Two ways: A) copy this folder (files included) — fastest. B) start from scratch — every file is printed in full below, paste and run.

⁠A) From this folder

cp .env.example .env               # set POSTGRES_PASSWORD
cp keys.example.json keys.json     # real apiKeys
cp budgets.example.json budgets.json
mkdir -p data
docker compose up -d
curl http://localhost:8787/health  # {"ok":true,...}

⁠B) From scratch

Create the files below side by side, then mkdir -p data && docker compose up -d.

⁠docker-compose.yml
name: my-aigateway

services:
  gateway:
    image: kienxuandaoit/my_aigateway:${GW_IMAGE_TAG:-latest}
    ports:
      # GW_HOST=127.0.0.1 for loopback-only (TLS proxy on same host).
      # Container side stays 8787 (app PORT default; change both if overridden).
      - "${GW_HOST:-0.0.0.0}:${GW_PORT:-8787}:8787"
    env_file:
      - path: .env
        required: false
    environment:
      # compose network hostname. overrides GW_DB from .env (usually a host sqlite path).
      GW_DB: postgres://postgres:${POSTGRES_PASSWORD:-secret}@db:5432/gw
    volumes:
      - ./keys.json:/app/keys.json:ro
      # presets + pricing catalog are standalone files (not baked into the
      # image): edit here, restart to apply. regen catalog: ./gen-models.sh
      - ./presets.json:/app/presets.json:ro
      - ./models.json:/app/models.json:ro
      # must exist on the host before `up`: Docker binds a missing path as a
      # *directory*, which the app refuses to read. `cp budgets.example.json budgets.json`.
      - ./budgets.json:/app/budgets.json:ro
      - ./data:/app/data
    restart: unless-stopped
    healthcheck:
      test: ["CMD-SHELL", "wget -qO- http://localhost:8787/health"]
      interval: 10s
      timeout: 3s
      retries: 3
      start_period: 10s
    depends_on:
      db:
        condition: service_healthy
  db:
    image: postgres:16-alpine
    environment:
      POSTGRES_DB: gw
      POSTGRES_PASSWORD: ${POSTGRES_PASSWORD:-secret}
    volumes:
      - pgdata:/var/lib/postgresql/data
    healthcheck:
      test: ["CMD-SHELL", "pg_isready -U postgres"]
      interval: 2s
      timeout: 3s
      retries: 20
    # not published — host already has :5432. query via `docker compose exec db psql`

volumes:
  pgdata:
⁠.env

Image tag + db password. Gateway GW_* vars go here too.

# cp .env.example .env
# Image tag to run. Must match a tag pushed to Docker Hub
# (./scripts/docker-push.sh [tag] pushes package.json version by default).
GW_IMAGE_TAG=0.1.0
# Postgres password for the bundled db service. Change it.
POSTGRES_PASSWORD=secret
# Gateway GW_* vars go here too (compose env_file passes them through).
# Full list with defaults: ../.env.example
# GW_UI_BASIC=ui:REPLACE-ME
# GW_ROLLUP=1h,24h
# GW_RETENTION=30d
⁠keys.json

Virtual keys (left side = Bearer token your clients send) → upstream routes. One entry per sdk. Replace every replace-me with a real key.

{
  "gw-openai-completions-replace-me": {
    "team": "core",
    "member": "pool",
    "routes": [
      {
        "sdk": "openai-completions",
        "provider": "google",
        "apiKey": "AIza-replace-me",
        "model": "gemini-3.5-flash-lite"
      }
    ]
  },
  "gw-openai-responses-replace-me": {
    "team": "core",
    "member": "pool",
    "routes": [
      {
        "sdk": "openai-responses",
        "provider": "bedrock-runtime.us-east-1",
        "apiKey": "replace-me",
        "model": "global.openai.gpt-5.6-sol"
      }
    ]
  },
  "gw-google-replace-me": {
    "team": "core",
    "member": "pool",
    "routes": [
      {
        "sdk": "google-generative-ai",
        "provider": "google",
        "apiKey": "AIza-replace-me",
        "model": "gemini-3.5-flash-lite"
      }
    ]
  },
  "gw-anthropic-messages-replace-me": {
    "team": "core",
    "member": "pool",
    "routes": [
      {
        "sdk": "anthropic-messages",
        "provider": "bedrock-runtime.us-east-1",
        "apiKey": "replace-me",
        "model": "anthropic.claude-sonnet-5"
      }
    ]
  },
  "gw-embeddings-replace-me": {
    "team": "core",
    "member": "pool",
    "routes": [
      {
        "sdk": "openai-embeddings",
        "provider": "ollama",
        "apiKey": "replace-me",
        "model": "nomic-embed-text"
      }
    ]
  }
}
⁠presets.json

Provider routing (hosts + per-wire paths). Rarely edited; add a block to point a provider name at another host. Missing file = loud boot error.

{
  "google": {
    "baseURL": "https://generativelanguage.googleapis.com",
    "wires": {
      "openai-completions": { "path": "v1beta/openai/chat/completions", "exclude": ["store"] },
      "google-generative-ai": { "path": "v1beta" }
    }
  },
  "openrouter": {
    "baseURL": "https://openrouter.ai/api",
    "wires": { "openai-completions": {} }
  },
  "logfare": {
    "baseURL": "https://logfare.ai",
    "wires": { "openai-completions": {} }
  },
  "poolside": {
    "baseURL": "https://inference.poolside.ai",
    "wires": { "openai-completions": {} }
  },
  "ollama": {
    "baseURL": "http://127.0.0.1:11434",
    "wires": {
      "openai-completions": {},
      "openai-embeddings": { "path": "v1/embeddings" }
    }
  },
  "opencode": {
    "baseURL": "https://opencode.ai/inference",
    "wires": {
      "openai-completions": { "path": "openai/v1/chat/completions" },
      "google-generative-ai": { "path": "google/v1beta" }
    }
  },
  "openai": {
    "baseURL": "https://api.openai.com",
    "wires": {
      "openai-completions": {},
      "openai-responses": {},
      "openai-embeddings": {}
    }
  },
  "anthropic": {
    "baseURL": "https://api.anthropic.com",
    "wires": { "anthropic-messages": {} }
  },
  "deepseek": {
    "baseURL": "https://api.deepseek.com",
    "wires": { "openai-completions": {} }
  },
  "xai": {
    "baseURL": "https://api.x.ai",
    "wires": { "openai-completions": {} }
  },
  "mistral": {
    "baseURL": "https://api.mistral.ai",
    "wires": { "openai-completions": {} }
  },
  "groq": {
    "baseURL": "https://api.groq.com/openai",
    "wires": { "openai-completions": {} }
  },
  "bedrock-runtime": {
    "baseURL": "https://bedrock-runtime.{region}.amazonaws.com",
    "wires": {
      "openai-completions": { "path": "openai/v1/chat/completions" },
      "openai-responses": { "path": "openai/v1/responses" },
      "anthropic-messages": { "path": "anthropic/v1/messages" }
    }
  },
  "bedrock-mantle": {
    "baseURL": "https://bedrock-mantle.{region}.api.aws",
    "wires": {
      "openai-completions": {},
      "openai-responses": { "path": "v1/responses" },
      "anthropic-messages": { "path": "anthropic/v1/messages" }
    }
  }
}
⁠models.json

Price catalog per serving path (provider/model, exactly as on the wire). Missing file = every est_cost null, spend caps silently never fire. Don't hand-write new pins — add them to keys.json, then regen:

cd docker                         # gen-models.sh lives beside keys.json
./gen-models.sh                   # tag from .env GW_IMAGE_TAG
./gen-models.sh 0.1.0             # ...or pick a tag explicitly

What it needs: keys.json with real pins (only provider/model are read, apiKeys unused), an existing models.json file (from-scratch: echo '{}' > models.json first — Docker binds a missing path as a directory), and network to models.dev. Host needs no node: fill runs as dist/models.js --fill inside the gateway image. Output per pin: fill (new entry), fill meta, refresh, manual-diff (hand-verified price kept, verify drift yourself), unmapped (no catalog match — add by hand, exit 1). Then docker compose restart gateway to apply.

Starter content:

{
  "google/gemini-3.5-flash-lite": {
    "input": 0.3,
    "output": 2.5,
    "cache_read": 0.03,
    "cache_write": null,
    "tiers": [],
    "meta": {
      "context_length": 1048576,
      "max_completion_tokens": 65536,
      "input_modalities": [
        "text",
        "image",
        "video",
        "audio",
        "pdf"
      ],
      "output_modalities": [
        "text"
      ]
    },
    "source": "models.dev",
    "ref": "google/gemini-3.5-flash-lite",
    "updatedAt": "2026-09-28T07:00:43.195Z"
  },
  "opencode/gemini-3.5-flash-lite": {
    "input": 0.3,
    "output": 2.5,
    "cache_read": 0.03,
    "cache_write": null,
    "tiers": [],
    "meta": {
      "context_length": 1048576,
      "max_completion_tokens": 65536,
      "input_modalities": [
        "text",
        "image",
        "video",
        "audio",
        "pdf"
      ],
      "output_modalities": [
        "text"
      ]
    },
    "source": "models.dev",
    "ref": "opencode/gemini-3.5-flash-lite",
    "updatedAt": "2026-09-28T07:00:43.195Z"
  },
  "opencode/muse-spark-1.3": {
    "input": 1.25,
    "output": 4.25,
    "cache_read": 0.15,
    "cache_write": null,
    "tiers": [],
    "meta": {
      "context_length": 1048576,
      "max_completion_tokens": 131072,
      "input_modalities": [
        "text",
        "image",
        "video",
        "pdf",
        "audio"
      ],
      "output_modalities": [
        "text"
      ]
    },
    "source": "models.dev",
    "ref": "opencode/muse-spark-1.3",
    "updatedAt": "2026-09-28T12:44:51.912Z"
  },
  "deepseek/deepseek-flash": {
    "input": 0.15,
    "output": 0.6,
    "cache_read": 0.003,
    "cache_write": null,
    "tiers": [],
    "meta": {
      "context_length": 1000000,
      "max_completion_tokens": 393216,
      "input_modalities": [
        "text",
        "image"
      ],
      "output_modalities": [
        "text"
      ]
    },
    "source": "manual",
    "ref": "https://api-docs.deepseek.com/quick_start/pricing (off-peak base, peak 01-04+06-10 UTC Mon-Fri x2)",
    "schedule": {
      "peakMult": 2,
      "days": [
        1,
        2,
        3,
        4,
        5
      ],
      "hours": [
        [
          1,
          4
        ],
        [
          6,
          10
        ]
      ]
    },
    "updatedAt": "2026-09-28T14:00:00.000Z"
  },
  "openrouter/openrouter/free": {
    "input": 0,
    "output": 0,
    "cache_read": null,
    "cache_write": null,
    "tiers": [],
    "meta": {
      "context_length": 200000,
      "max_completion_tokens": 8000,
      "input_modalities": [
        "text",
        "image"
      ],
      "output_modalities": [
        "text"
      ]
    },
    "source": "models.dev",
    "ref": "openrouter/openrouter/free",
    "updatedAt": "2026-09-30T04:06:49.384Z"
  },
  "opencode/big-pickle": {
    "input": 0,
    "output": 0,
    "cache_read": 0,
    "cache_write": 0,
    "tiers": [],
    "meta": {
      "context_length": 200000,
      "max_completion_tokens": 32000,
      "input_modalities": [
        "text"
      ],
      "output_modalities": [
        "text"
      ]
    },
    "source": "models.dev",
    "ref": "opencode/big-pickle",
    "updatedAt": "2026-09-30T04:06:49.386Z"
  },
  "bedrock-runtime.us-east-1/global.openai.gpt-5.6-sol": {
    "input": 4,
    "output": 20,
    "cache_read": 0.4,
    "cache_write": 5,
    "tiers": [
      {
        "size": 272000,
        "input": 8,
        "output": 30,
        "cache_read": 0.8
      }
    ],
    "meta": {
      "context_length": 1050000,
      "max_completion_tokens": 128000,
      "input_modalities": [
        "text",
        "image"
      ],
      "output_modalities": [
        "text"
      ]
    },
    "source": "models.dev",
    "ref": "amazon-bedrock/global.openai.gpt-5.6-sol",
    "updatedAt": "2026-09-30T04:06:49.387Z"
  },
  "bedrock-mantle.us-east-1/openai.gpt-oss-120b": {
    "input": 0.15,
    "output": 0.6,
    "cache_read": null,
    "cache_write": null,
    "tiers": [],
    "meta": {
      "context_length": 131072,
      "max_completion_tokens": 131072,
      "input_modalities": [
        "text"
      ],
      "output_modalities": [
        "text"
      ]
    },
    "source": "models.dev",
    "ref": "amazon-bedrock/openai.gpt-oss-120b",
    "updatedAt": "2026-09-30T04:06:49.388Z"
  },
  "bedrock-runtime.us-east-1/anthropic.claude-sonnet-5": {
    "input": 2,
    "output": 10,
    "cache_read": 0.2,
    "cache_write": 2.5,
    "tiers": [],
    "meta": {
      "context_length": 1000000,
      "max_completion_tokens": 128000,
      "input_modalities": [
        "text",
        "image",
        "pdf"
      ],
      "output_modalities": [
        "text"
      ]
    },
    "source": "models.dev",
    "ref": "amazon-bedrock/anthropic.claude-sonnet-5",
    "updatedAt": "2026-09-30T04:06:49.388Z"
  },
  "bedrock-mantle.eu-central-1/anthropic.claude-sonnet-5": {
    "input": 2,
    "output": 10,
    "cache_read": 0.2,
    "cache_write": 2.5,
    "tiers": [],
    "meta": {
      "context_length": 1000000,
      "max_completion_tokens": 128000,
      "input_modalities": [
        "text",
        "image",
        "pdf"
      ],
      "output_modalities": [
        "text"
      ]
    },
    "source": "models.dev",
    "ref": "amazon-bedrock/anthropic.claude-sonnet-5",
    "updatedAt": "2026-09-30T04:06:49.389Z"
  },
  "ollama/nomic-embed-text": {
    "input": 0.05,
    "output": 0,
    "cache_read": null,
    "cache_write": null,
    "tiers": [],
    "meta": {
      "context_length": 8192,
      "max_completion_tokens": 768,
      "input_modalities": [
        "text"
      ],
      "output_modalities": [
        "text"
      ]
    },
    "source": "models.dev",
    "ref": "nomic-embed-text",
    "updatedAt": "2026-09-30T04:06:49.390Z"
  }
}
⁠budgets.json

Symbolic spend caps (reference estimates, not bills). Empty team/member/model = wildcard. Must exist before up.

{
  "_comment": "cp budgets.example.json budgets.json. Symbolic caps on reference estimates (est_cost), not bills. Empty team/member/model = wildcard. alerts fire via rollup tick (log) + /ui pill. \"enforce\": true additionally rejects requests past the cap with 429; \"concurrency\" (needs enforce) caps in-flight requests for the scope.",
  "budgets": [
    { "team": "core", "member": "pool", "window": "24h", "usd": 10, "enforce": true, "concurrency": 4 },
    { "team": "core", "member": "pool", "model": "flash", "window": "7d", "usd": 20 }
  ]
}

⁠Ops

  • Verify: curl http://localhost:8787/health → {"ok":true,...}.
  • Chat call: curl -X POST http://localhost:8787/v1/chat/completions -H "Authorization: Bearer <your-token>" -H "Content-Type: application/json" -d '{"model":"<pin>","messages":[{"role":"user","content":"ping"}]}'.
  • Query spend: docker compose exec db psql -U postgres -d gw.
  • Update image: set GW_IMAGE_TAG in .env, docker compose up -d.
  • Plain HTTP only — TLS-terminating reverse proxy (Caddy/nginx) in front beyond localhost. Same-host proxy: GW_HOST=127.0.0.1.
  • No keys/.env baked into the image. Secrets stay mounts/env.

Tag summary

Content type

Image

Digest

sha256:c52f69ca1…

Size

59 MB

Last updated

3 days ago

docker pull kienxuandaoit/my_aigateway