sglang/docs/basic_usage/openai_api_embeddings.ipynb

{
 "cells": [
  {
   "cell_type": "markdown",
   "metadata": {},
   "source": [
    "# OpenAI APIs - Embedding\n",
    "\n",
    "SGLang provides OpenAI-compatible APIs to enable a smooth transition from OpenAI services to self-hosted local models.\n",
    "A complete reference for the API is available in the [OpenAI API Reference](https://platform.openai.com/docs/guides/embeddings).\n",
    "\n",
    "This tutorial covers the embedding APIs for embedding models. For a list of the supported models see the [corresponding overview page](../supported_models/embedding_models.md)\n"
   ]
  },
  {
   "cell_type": "markdown",
   "metadata": {},
   "source": [
    "## Launch A Server\n",
    "\n",
    "Launch the server in your terminal and wait for it to initialize. Remember to add `--is-embedding` to the command."
   ]
  },
  {
   "cell_type": "code",
   "execution_count": null,
   "metadata": {},
   "outputs": [],
   "source": [
    "from sglang.test.doc_patch import launch_server_cmd\n",
    "from sglang.utils import wait_for_server, print_highlight, terminate_process\n",
    "\n",
    "embedding_process, port = launch_server_cmd(\n",
    "    \"\"\"\n",
    "python3 -m sglang.launch_server --model-path Alibaba-NLP/gte-Qwen2-1.5B-instruct \\\n",
    "    --host 0.0.0.0 --is-embedding --log-level warning\n",
    "\"\"\"\n",
    ")\n",
    "\n",
    "wait_for_server(f\"http://localhost:{port}\")"
   ]
  },
  {
   "cell_type": "markdown",
   "metadata": {},
   "source": [
    "## Using cURL"
   ]
  },
  {
   "cell_type": "code",
   "execution_count": null,
   "metadata": {},
   "outputs": [],
   "source": [
    "import subprocess, json\n",
    "\n",
    "text = \"Once upon a time\"\n",
    "\n",
    "curl_text = f\"\"\"curl -s http://localhost:{port}/v1/embeddings \\\n",
    "  -H \"Content-Type: application/json\" \\\n",
    "  -d '{{\"model\": \"Alibaba-NLP/gte-Qwen2-1.5B-instruct\", \"input\": \"{text}\"}}'\"\"\"\n",
    "\n",
    "result = subprocess.check_output(curl_text, shell=True)\n",
    "\n",
    "print(result)\n",
    "\n",
    "text_embedding = json.loads(result)[\"data\"][0][\"embedding\"]\n",
    "\n",
    "print_highlight(f\"Text embedding (first 10): {text_embedding[:10]}\")"
   ]
  },
  {
   "cell_type": "markdown",
   "metadata": {},
   "source": [
    "## Using Python Requests"
   ]
  },
  {
   "cell_type": "code",
   "execution_count": null,
   "metadata": {},
   "outputs": [],
   "source": [
    "import requests\n",
    "\n",
    "text = \"Once upon a time\"\n",
    "\n",
    "response = requests.post(\n",
    "    f\"http://localhost:{port}/v1/embeddings\",\n",
    "    json={\"model\": \"Alibaba-NLP/gte-Qwen2-1.5B-instruct\", \"input\": text},\n",
    ")\n",
    "\n",
    "text_embedding = response.json()[\"data\"][0][\"embedding\"]\n",
    "\n",
    "print_highlight(f\"Text embedding (first 10): {text_embedding[:10]}\")"
   ]
  },
  {
   "cell_type": "markdown",
   "metadata": {},
   "source": [
    "## Using OpenAI Python Client"
   ]
  },
  {
   "cell_type": "code",
   "execution_count": null,
   "metadata": {},
   "outputs": [],
   "source": [
    "import openai\n",
    "\n",
    "client = openai.Client(base_url=f\"http://127.0.0.1:{port}/v1\", api_key=\"None\")\n",
    "\n",
    "# Text embedding example\n",
    "response = client.embeddings.create(\n",
    "    model=\"Alibaba-NLP/gte-Qwen2-1.5B-instruct\",\n",
    "    input=text,\n",
    ")\n",
    "\n",
    "embedding = response.data[0].embedding[:10]\n",
    "print_highlight(f\"Text embedding (first 10): {embedding}\")"
   ]
  },
  {
   "cell_type": "markdown",
   "metadata": {},
   "source": [
    "## Using Input IDs\n",
    "\n",
    "SGLang also supports `input_ids` as input to get the embedding."
   ]
  },
  {
   "cell_type": "code",
   "execution_count": null,
   "metadata": {},
   "outputs": [],
   "source": [
    "import json\n",
    "import os\n",
    "from transformers import AutoTokenizer\n",
    "\n",
    "os.environ[\"TOKENIZERS_PARALLELISM\"] = \"false\"\n",
    "\n",
    "tokenizer = AutoTokenizer.from_pretrained(\"Alibaba-NLP/gte-Qwen2-1.5B-instruct\")\n",
    "input_ids = tokenizer.encode(text)\n",
    "\n",
    "curl_ids = f\"\"\"curl -s http://localhost:{port}/v1/embeddings \\\n",
    "  -H \"Content-Type: application/json\" \\\n",
    "  -d '{{\"model\": \"Alibaba-NLP/gte-Qwen2-1.5B-instruct\", \"input\": {json.dumps(input_ids)}}}'\"\"\"\n",
    "\n",
    "input_ids_embedding = json.loads(subprocess.check_output(curl_ids, shell=True))[\"data\"][\n",
    "    0\n",
    "][\"embedding\"]\n",
    "\n",
    "print_highlight(f\"Input IDs embedding (first 10): {input_ids_embedding[:10]}\")"
   ]
  },
  {
   "cell_type": "code",
   "execution_count": null,
   "metadata": {},
   "outputs": [],
   "source": [
    "terminate_process(embedding_process)"
   ]
  },
  {
   "cell_type": "markdown",
   "metadata": {},
   "source": [
    "## Multi-Modal Embedding Model\n",
    "Please refer to [Multi-Modal Embedding Model](../supported_models/embedding_models.md)"
   ]
  }
 ],
 "metadata": {
  "language_info": {
   "codemirror_mode": {
    "name": "ipython",
    "version": 3
   },
   "file_extension": ".py",
   "mimetype": "text/x-python",
   "name": "python",
   "nbconvert_exporter": "python",
   "pygments_lexer": "ipython3"
  }
 },
 "nbformat": 4,
 "nbformat_minor": 2
}
Add support for ipynb (#1786) 2024-10-25 20:48:35 -07:00			`{`
			`"cells": [`
			`{`
			`"cell_type": "markdown",`
			`"metadata": {},`
			`"source": [`
Native api (#1886) Co-authored-by: Chayenne <zhaochenyang@g.ucla.edu> 2024-11-02 01:02:17 -07:00			`"# OpenAI APIs - Embedding\n",`
Add openAI compatible API (#1810) Co-authored-by: Chayenne <zhaochenyang@g.ucla.edu> 2024-10-27 10:51:42 -07:00			`"\n",`
Fix docs (#1889) 2024-11-02 11:46:00 -07:00			`"SGLang provides OpenAI-compatible APIs to enable a smooth transition from OpenAI services to self-hosted local models.\n",`
			`"A complete reference for the API is available in the [OpenAI API Reference](https://platform.openai.com/docs/guides/embeddings).\n",`
Add openAI compatible API (#1810) Co-authored-by: Chayenne <zhaochenyang@g.ucla.edu> 2024-10-27 10:51:42 -07:00			`"\n",`
Refactor the docs (#9031) 2025-08-10 19:49:45 -07:00			`"This tutorial covers the embedding APIs for embedding models. For a list of the supported models see the [corresponding overview page](../supported_models/embedding_models.md)\n"`
Add support for ipynb (#1786) 2024-10-25 20:48:35 -07:00			`]`
			`},`
			`{`
			`"cell_type": "markdown",`
			`"metadata": {},`
			`"source": [`
Add openAI compatible API (#1810) Co-authored-by: Chayenne <zhaochenyang@g.ucla.edu> 2024-10-27 10:51:42 -07:00			`"## Launch A Server\n",`
			`"\n",`
Docs: Refactor Contribution Guide (#2690) 2024-12-31 22:11:00 +00:00			"Launch the server in your terminal and wait for it to initialize. Remember to add `--is-embedding` to the command."
Add support for ipynb (#1786) 2024-10-25 20:48:35 -07:00			`]`
			`},`
			`{`
			`"cell_type": "code",`
add native api docs (#1883) Co-authored-by: Chayenne <zhaochenyang@g.ucla.edu> 2024-11-02 00:17:30 -07:00			`"execution_count": null,`
feat(pre-commit): trim unnecessary notebook metadata from git history (#2127) 2024-11-23 05:04:51 +08:00			`"metadata": {},`
add native api docs (#1883) Co-authored-by: Chayenne <zhaochenyang@g.ucla.edu> 2024-11-02 00:17:30 -07:00			`"outputs": [],`
Add support for ipynb (#1786) 2024-10-25 20:48:35 -07:00			`"source": [`
Refactor the docs (#9031) 2025-08-10 19:49:45 -07:00			`"from sglang.test.doc_patch import launch_server_cmd\n",`
[CI] Improve Docs CI Efficiency (#3587) Co-authored-by: zhaochenyang20 <zhaochen20@outlook.com> 2025-02-15 03:57:00 +00:00			`"from sglang.utils import wait_for_server, print_highlight, terminate_process\n",`
Simplify our docs with complicated functions into utils (#1807) Co-authored-by: Chayenne <zhaochenyang@ucla.edu> 2024-10-26 10:44:11 -07:00			`"\n",`
[CI] Improve Docs CI Efficiency (#3587) Co-authored-by: zhaochenyang20 <zhaochen20@outlook.com> 2025-02-15 03:57:00 +00:00			`"embedding_process, port = launch_server_cmd(\n",`
Add openAI compatible API (#1810) Co-authored-by: Chayenne <zhaochenyang@g.ucla.edu> 2024-10-27 10:51:42 -07:00			`" \"\"\"\n",`
smaller and non gated models for docs (#5378) 2025-04-21 02:38:25 +02:00			`"python3 -m sglang.launch_server --model-path Alibaba-NLP/gte-Qwen2-1.5B-instruct \\\n",`
[Doc] Fix SGLang tool parser doc (#9886) 2025-09-04 09:52:53 -04:00			`" --host 0.0.0.0 --is-embedding --log-level warning\n",`
Add openAI compatible API (#1810) Co-authored-by: Chayenne <zhaochenyang@g.ucla.edu> 2024-10-27 10:51:42 -07:00			`"\"\"\"\n",`
			`")\n",`
Add support for ipynb (#1786) 2024-10-25 20:48:35 -07:00			`"\n",`
[CI] Improve Docs CI Efficiency (#3587) Co-authored-by: zhaochenyang20 <zhaochen20@outlook.com> 2025-02-15 03:57:00 +00:00			`"wait_for_server(f\"http://localhost:{port}\")"`
Add support for ipynb (#1786) 2024-10-25 20:48:35 -07:00			`]`
			`},`
			`{`
			`"cell_type": "markdown",`
			`"metadata": {},`
			`"source": [`
Native api (#1886) Co-authored-by: Chayenne <zhaochenyang@g.ucla.edu> 2024-11-02 01:02:17 -07:00			`"## Using cURL"`
Add support for ipynb (#1786) 2024-10-25 20:48:35 -07:00			`]`
			`},`
			`{`
			`"cell_type": "code",`
add native api docs (#1883) Co-authored-by: Chayenne <zhaochenyang@g.ucla.edu> 2024-11-02 00:17:30 -07:00			`"execution_count": null,`
feat(pre-commit): trim unnecessary notebook metadata from git history (#2127) 2024-11-23 05:04:51 +08:00			`"metadata": {},`
add native api docs (#1883) Co-authored-by: Chayenne <zhaochenyang@g.ucla.edu> 2024-11-02 00:17:30 -07:00			`"outputs": [],`
Add support for ipynb (#1786) 2024-10-25 20:48:35 -07:00			`"source": [`
Add openAI compatible API (#1810) Co-authored-by: Chayenne <zhaochenyang@g.ucla.edu> 2024-10-27 10:51:42 -07:00			`"import subprocess, json\n",`
			`"\n",`
			`"text = \"Once upon a time\"\n",`
Add support for ipynb (#1786) 2024-10-25 20:48:35 -07:00			`"\n",`
[CI] Improve Docs CI Efficiency (#3587) Co-authored-by: zhaochenyang20 <zhaochen20@outlook.com> 2025-02-15 03:57:00 +00:00			`"curl_text = f\"\"\"curl -s http://localhost:{port}/v1/embeddings \\\n",`
feat(oai refactor): Replace `openai_api` with `entrypoints/openai` (#7351) Co-authored-by: Jin Pan <jpan236@wisc.edu> 2025-06-21 13:21:06 -07:00			`" -H \"Content-Type: application/json\" \\\n",`
smaller and non gated models for docs (#5378) 2025-04-21 02:38:25 +02:00			`" -d '{{\"model\": \"Alibaba-NLP/gte-Qwen2-1.5B-instruct\", \"input\": \"{text}\"}}'\"\"\"\n",`
Add openAI compatible API (#1810) Co-authored-by: Chayenne <zhaochenyang@g.ucla.edu> 2024-10-27 10:51:42 -07:00			`"\n",`
feat(oai refactor): Replace `openai_api` with `entrypoints/openai` (#7351) Co-authored-by: Jin Pan <jpan236@wisc.edu> 2025-06-21 13:21:06 -07:00			`"result = subprocess.check_output(curl_text, shell=True)\n",`
			`"\n",`
			`"print(result)\n",`
			`"\n",`
			`"text_embedding = json.loads(result)[\"data\"][0][\"embedding\"]\n",`
Add openAI compatible API (#1810) Co-authored-by: Chayenne <zhaochenyang@g.ucla.edu> 2024-10-27 10:51:42 -07:00			`"\n",`
Imporve openai api documents (#1827) Co-authored-by: Chayenne <zhaochenyang@g.ucla.edu> 2024-10-30 00:39:41 -07:00			`"print_highlight(f\"Text embedding (first 10): {text_embedding[:10]}\")"`
Add support for ipynb (#1786) 2024-10-25 20:48:35 -07:00			`]`
			`},`
			`{`
			`"cell_type": "markdown",`
			`"metadata": {},`
			`"source": [`
Fix docs (#1889) 2024-11-02 11:46:00 -07:00			`"## Using Python Requests"`
Native api (#1886) Co-authored-by: Chayenne <zhaochenyang@g.ucla.edu> 2024-11-02 01:02:17 -07:00			`]`
			`},`
			`{`
			`"cell_type": "code",`
			`"execution_count": null,`
feat(pre-commit): trim unnecessary notebook metadata from git history (#2127) 2024-11-23 05:04:51 +08:00			`"metadata": {},`
Native api (#1886) Co-authored-by: Chayenne <zhaochenyang@g.ucla.edu> 2024-11-02 01:02:17 -07:00			`"outputs": [],`
			`"source": [`
			`"import requests\n",`
			`"\n",`
			`"text = \"Once upon a time\"\n",`
			`"\n",`
			`"response = requests.post(\n",`
[CI] Improve Docs CI Efficiency (#3587) Co-authored-by: zhaochenyang20 <zhaochen20@outlook.com> 2025-02-15 03:57:00 +00:00			`" f\"http://localhost:{port}/v1/embeddings\",\n",`
smaller and non gated models for docs (#5378) 2025-04-21 02:38:25 +02:00			`" json={\"model\": \"Alibaba-NLP/gte-Qwen2-1.5B-instruct\", \"input\": text},\n",`
Native api (#1886) Co-authored-by: Chayenne <zhaochenyang@g.ucla.edu> 2024-11-02 01:02:17 -07:00			`")\n",`
			`"\n",`
			`"text_embedding = response.json()[\"data\"][0][\"embedding\"]\n",`
			`"\n",`
			`"print_highlight(f\"Text embedding (first 10): {text_embedding[:10]}\")"`
			`]`
			`},`
			`{`
			`"cell_type": "markdown",`
			`"metadata": {},`
			`"source": [`
			`"## Using OpenAI Python Client"`
Add support for ipynb (#1786) 2024-10-25 20:48:35 -07:00			`]`
			`},`
			`{`
			`"cell_type": "code",`
add native api docs (#1883) Co-authored-by: Chayenne <zhaochenyang@g.ucla.edu> 2024-11-02 00:17:30 -07:00			`"execution_count": null,`
feat(pre-commit): trim unnecessary notebook metadata from git history (#2127) 2024-11-23 05:04:51 +08:00			`"metadata": {},`
add native api docs (#1883) Co-authored-by: Chayenne <zhaochenyang@g.ucla.edu> 2024-11-02 00:17:30 -07:00			`"outputs": [],`
Add support for ipynb (#1786) 2024-10-25 20:48:35 -07:00			`"source": [`
			`"import openai\n",`
			`"\n",`
[CI] Improve Docs CI Efficiency (#3587) Co-authored-by: zhaochenyang20 <zhaochen20@outlook.com> 2025-02-15 03:57:00 +00:00			`"client = openai.Client(base_url=f\"http://127.0.0.1:{port}/v1\", api_key=\"None\")\n",`
Add support for ipynb (#1786) 2024-10-25 20:48:35 -07:00			`"\n",`
			`"# Text embedding example\n",`
			`"response = client.embeddings.create(\n",`
smaller and non gated models for docs (#5378) 2025-04-21 02:38:25 +02:00			`" model=\"Alibaba-NLP/gte-Qwen2-1.5B-instruct\",\n",`
Add openAI compatible API (#1810) Co-authored-by: Chayenne <zhaochenyang@g.ucla.edu> 2024-10-27 10:51:42 -07:00			`" input=text,\n",`
Add support for ipynb (#1786) 2024-10-25 20:48:35 -07:00			`")\n",`
			`"\n",`
			`"embedding = response.data[0].embedding[:10]\n",`
Imporve openai api documents (#1827) Co-authored-by: Chayenne <zhaochenyang@g.ucla.edu> 2024-10-30 00:39:41 -07:00			`"print_highlight(f\"Text embedding (first 10): {embedding}\")"`
Add openAI compatible API (#1810) Co-authored-by: Chayenne <zhaochenyang@g.ucla.edu> 2024-10-27 10:51:42 -07:00			`]`
			`},`
			`{`
			`"cell_type": "markdown",`
			`"metadata": {},`
			`"source": [`
			`"## Using Input IDs\n",`
			`"\n",`
			"SGLang also supports `input_ids` as input to get the embedding."
			`]`
			`},`
			`{`
			`"cell_type": "code",`
add native api docs (#1883) Co-authored-by: Chayenne <zhaochenyang@g.ucla.edu> 2024-11-02 00:17:30 -07:00			`"execution_count": null,`
feat(pre-commit): trim unnecessary notebook metadata from git history (#2127) 2024-11-23 05:04:51 +08:00			`"metadata": {},`
add native api docs (#1883) Co-authored-by: Chayenne <zhaochenyang@g.ucla.edu> 2024-11-02 00:17:30 -07:00			`"outputs": [],`
Add openAI compatible API (#1810) Co-authored-by: Chayenne <zhaochenyang@g.ucla.edu> 2024-10-27 10:51:42 -07:00			`"source": [`
			`"import json\n",`
			`"import os\n",`
			`"from transformers import AutoTokenizer\n",`
			`"\n",`
			`"os.environ[\"TOKENIZERS_PARALLELISM\"] = \"false\"\n",`
			`"\n",`
smaller and non gated models for docs (#5378) 2025-04-21 02:38:25 +02:00			`"tokenizer = AutoTokenizer.from_pretrained(\"Alibaba-NLP/gte-Qwen2-1.5B-instruct\")\n",`
Add openAI compatible API (#1810) Co-authored-by: Chayenne <zhaochenyang@g.ucla.edu> 2024-10-27 10:51:42 -07:00			`"input_ids = tokenizer.encode(text)\n",`
			`"\n",`
[CI] Improve Docs CI Efficiency (#3587) Co-authored-by: zhaochenyang20 <zhaochen20@outlook.com> 2025-02-15 03:57:00 +00:00			`"curl_ids = f\"\"\"curl -s http://localhost:{port}/v1/embeddings \\\n",`
feat(oai refactor): Replace `openai_api` with `entrypoints/openai` (#7351) Co-authored-by: Jin Pan <jpan236@wisc.edu> 2025-06-21 13:21:06 -07:00			`" -H \"Content-Type: application/json\" \\\n",`
smaller and non gated models for docs (#5378) 2025-04-21 02:38:25 +02:00			`" -d '{{\"model\": \"Alibaba-NLP/gte-Qwen2-1.5B-instruct\", \"input\": {json.dumps(input_ids)}}}'\"\"\"\n",`
Add openAI compatible API (#1810) Co-authored-by: Chayenne <zhaochenyang@g.ucla.edu> 2024-10-27 10:51:42 -07:00			`"\n",`
			`"input_ids_embedding = json.loads(subprocess.check_output(curl_ids, shell=True))[\"data\"][\n",`
			`" 0\n",`
			`"][\"embedding\"]\n",`
			`"\n",`
Imporve openai api documents (#1827) Co-authored-by: Chayenne <zhaochenyang@g.ucla.edu> 2024-10-30 00:39:41 -07:00			`"print_highlight(f\"Input IDs embedding (first 10): {input_ids_embedding[:10]}\")"`
Add support for ipynb (#1786) 2024-10-25 20:48:35 -07:00			`]`
Simplify our docs with complicated functions into utils (#1807) Co-authored-by: Chayenne <zhaochenyang@ucla.edu> 2024-10-26 10:44:11 -07:00			`},`
			`{`
			`"cell_type": "code",`
feat(pre-commit): trim unnecessary notebook metadata from git history (#2127) 2024-11-23 05:04:51 +08:00			`"execution_count": null,`
			`"metadata": {},`
change file tree (#1859) Co-authored-by: Chayenne <zhaochenyang@g.ucla.edu> 2024-10-31 20:10:16 -07:00			`"outputs": [],`
Simplify our docs with complicated functions into utils (#1807) Co-authored-by: Chayenne <zhaochenyang@ucla.edu> 2024-10-26 10:44:11 -07:00			`"source": [`
[Docs]: Fix Multi-User Port Allocation Conflicts (#3601) Co-authored-by: zhaochenyang20 <zhaochen20@outlook.com> Co-authored-by: simveit <simp.veitner@gmail.com> 2025-02-19 19:15:44 +00:00			`"terminate_process(embedding_process)"`
Simplify our docs with complicated functions into utils (#1807) Co-authored-by: Chayenne <zhaochenyang@ucla.edu> 2024-10-26 10:44:11 -07:00			`]`
doc: fix the erroneous documents and example codes about Alibaba-NLP/gme-Qwen2-VL-2B-Instruct (#6199) 2025-05-11 23:22:11 +08:00			`},`
			`{`
			`"cell_type": "markdown",`
			`"metadata": {},`
			`"source": [`
			`"## Multi-Modal Embedding Model\n",`
			`"Please refer to [Multi-Modal Embedding Model](../supported_models/embedding_models.md)"`
			`]`
Add support for ipynb (#1786) 2024-10-25 20:48:35 -07:00			`}`
			`],`
			`"metadata": {`
			`"language_info": {`
			`"codemirror_mode": {`
			`"name": "ipython",`
			`"version": 3`
			`},`
			`"file_extension": ".py",`
			`"mimetype": "text/x-python",`
			`"name": "python",`
			`"nbconvert_exporter": "python",`
feat(pre-commit): trim unnecessary notebook metadata from git history (#2127) 2024-11-23 05:04:51 +08:00			`"pygments_lexer": "ipython3"`
Add support for ipynb (#1786) 2024-10-25 20:48:35 -07:00			`}`
			`},`
			`"nbformat": 4,`
			`"nbformat_minor": 2`
			`}`