Compare commits
93 Commits
7ff076530a
...
main
| Author | SHA1 | Date | |
|---|---|---|---|
| 76dd476019 | |||
| f2c3445113 | |||
| 18ba470f4c | |||
| f3e90d2f17 | |||
| 18a2169a52 | |||
| 40f8c5ac05 | |||
| 860b711e2a | |||
| e932ef0915 | |||
| 1bc448f505 | |||
| 2d6c3dcb82 | |||
| 056cfde235 | |||
| eedefec4fb | |||
| 4b02041c1d | |||
| 2d0b2f4aae | |||
| 9dedf25480 | |||
| 5f85139bd4 | |||
| 6e4e952e20 | |||
| 6a3403a530 | |||
| a5b0402c9d | |||
| 719c593c82 | |||
| b534ceac58 | |||
| f595b53d70 | |||
| b1f0fd6a39 | |||
| 405e17ea3c | |||
| 00d34a4c9a | |||
| 7377cadffa | |||
| 6fbfba3505 | |||
| 99025e1c7e | |||
| 44ec299348 | |||
| 53914cfb8b | |||
| 80a549bdcd | |||
| d87daadc2f | |||
| 4aeaebdd3e | |||
| 2d3f982d61 | |||
| b12c6e6ec0 | |||
| 101ccb3553 | |||
| a13fb0dd23 | |||
| e46558bc28 | |||
| db43278cb7 | |||
| 0f49bed00e | |||
| 8d5e52fcaa | |||
| af4b7fcaa2 | |||
| cdebc8ac36 | |||
| 6f2fd6be9c | |||
| b5feb9702b | |||
| 549cc05603 | |||
| fd579f68ba | |||
| 0027c758d7 | |||
| 240fcc67ae | |||
| 41aa18e42f | |||
| d7bb408924 | |||
| 7ab9fe6e59 | |||
| 9151dcf576 | |||
| b009a99392 | |||
| aec7329be5 | |||
| 801aa8f0bb | |||
| fc86b4426b | |||
| 87aaef7f0a | |||
| edc6c5af84 | |||
| 80c45e6749 | |||
| 3b4a523b38 | |||
| 6574035969 | |||
| 1f807347ee | |||
| ba1d64f7a3 | |||
| ccc2140d76 | |||
| 0c6314d2f7 | |||
| 69f567a2f4 | |||
| 82c2a1dd81 | |||
| f97428219f | |||
| e94e83e91f | |||
| 1622a3d472 | |||
| 491e2ba7e5 | |||
| d5e5911031 | |||
| 24654096f1 | |||
| 1a62b156ef | |||
| e59ca145be | |||
| 234c59b0b4 | |||
| dbd934cdda | |||
| 1aa8d011e8 | |||
| a046930b43 | |||
| 6f18a26bd9 | |||
| f82de29267 | |||
| 1a61a7bb17 | |||
| 10925e742b | |||
| afea99ad2d | |||
| 3f2f2f6f7d | |||
| dfc7c25b5f | |||
| 47adb86693 | |||
| 757a29d7ee | |||
| eef84a2199 | |||
| d8a0076339 | |||
| bccafdfa54 | |||
| d84cdc00e5 |
@@ -0,0 +1,2 @@
|
||||
.idea/
|
||||
venv/
|
||||
+4
-2
@@ -1,9 +1,11 @@
|
||||
FROM docker.1ms.run/vllm/vllm-openai-rocm:latest
|
||||
FROM vllm/vllm-openai-rocm:nightly
|
||||
|
||||
WORKDIR /workspace
|
||||
|
||||
COPY requirements.txt /workspace/requirements.txt
|
||||
RUN pip install --no-cache-dir --retries 10 --timeout 180 -r /workspace/requirements.txt
|
||||
RUN pip install --no-cache-dir --retries 20 --timeout 600 -i https://pypi.tuna.tsinghua.edu.cn/simple -r /workspace/requirements.txt \
|
||||
|| pip install --no-cache-dir --retries 20 --timeout 600 -i https://mirrors.aliyun.com/pypi/simple -r /workspace/requirements.txt \
|
||||
|| pip install --no-cache-dir --retries 20 --timeout 600 -r /workspace/requirements.txt
|
||||
|
||||
COPY app /workspace/app
|
||||
|
||||
|
||||
@@ -15,8 +15,7 @@
|
||||
├── app
|
||||
│ ├── config.py
|
||||
│ ├── model_catalog.py
|
||||
│ ├── start_openai.py
|
||||
│ └── schemas.py
|
||||
│ └── start_openai.py
|
||||
├── .dockerignore
|
||||
├── config.json
|
||||
├── docker-compose.yml
|
||||
@@ -29,6 +28,9 @@
|
||||
项目只读取一个配置文件:`config.json`。
|
||||
|
||||
- `services.openai.host` / `services.openai.port`:OpenAI 协议服务监听地址与端口(默认 `0.0.0.0:8001`)
|
||||
- `public_model_name`:对外固定模型名,切换底层模型时可保持调用方参数不变
|
||||
- `default_enable_thinking`:服务端默认思考开关,默认 `false`(即调用方不传时也关闭)
|
||||
- `reasoning_enabled`:是否启用推理解析器参数注入,默认 `false`
|
||||
- `api_key`:OpenAI 接口访问密钥
|
||||
- `tensor_parallel_size`:张量并行数,双卡建议 `2`
|
||||
- `dtype`:推理精度,默认 `bfloat16`
|
||||
@@ -43,7 +45,7 @@
|
||||
- `models.default`:默认模型名
|
||||
- `models.selected`:当前生效模型名
|
||||
- `models.profiles`:模型配置集合
|
||||
- 每个模型必须包含:`local_path`,并建议补充 `ctx`、`max_num_seqs`、`max_tokens`、`dtype`、`quantization`
|
||||
- 每个模型必须包含:`local_path`,并建议补充 `ctx`、`max_num_seqs`、`max_tokens`、`dtype`、`quantization`、`reasoning_parser`
|
||||
|
||||
启动时会按以下优先级选模型:
|
||||
|
||||
@@ -62,34 +64,52 @@
|
||||
|
||||
## 部署步骤
|
||||
|
||||
1. 修改 `config.json` 中的 `models.selected` 与服务参数。
|
||||
1. 预拉取基础镜像(与官方文档一致):
|
||||
|
||||
2. 构建并启动容器:
|
||||
```bash
|
||||
docker pull docker.1ms.run/vllm/vllm-openai-rocm:latest
|
||||
```
|
||||
|
||||
2. 修改 `config.json` 中的 `models.selected` 与服务参数。
|
||||
|
||||
3. 构建并启动容器:
|
||||
|
||||
```bash
|
||||
docker compose up -d --build
|
||||
```
|
||||
|
||||
3. 验证 OpenAI 协议服务:
|
||||
4. 验证 OpenAI 协议服务:
|
||||
|
||||
```bash
|
||||
curl http://localhost:<services.openai.port>/v1/models
|
||||
```
|
||||
|
||||
当前 `docker-compose.yml` 已按官方运行参数适配:
|
||||
|
||||
- `--group-add=video` → `group_add: [video]`
|
||||
- `--ipc=host` → `ipc: host`
|
||||
- `--cap-add=SYS_PTRACE` → `cap_add: [SYS_PTRACE]`
|
||||
- `--security-opt seccomp=unconfined` → `security_opt: [seccomp=unconfined]`
|
||||
- `--device /dev/kfd` 与 `--device /dev/dri` → `devices`
|
||||
- `-e HF_HOME=/app/models` 在本项目等效为 `HF_HOME=/opt/model`
|
||||
|
||||
## OpenAI 协议示例(8001)
|
||||
|
||||
```bash
|
||||
curl -X POST "http://localhost:8001/v1/chat/completions" \
|
||||
-H "Content-Type: application/json" \
|
||||
-H "Authorization: Bearer <config.json中的api_key>" \
|
||||
-d "{\"model\":\"Qwen3.5-35B-A3B-GPTQ-Int4\",\"messages\":[{\"role\":\"user\",\"content\":\"你好,介绍一下你自己\"}],\"temperature\":0.7}"
|
||||
-d "{\"model\":\"Qwen_local_model\",\"messages\":[{\"role\":\"user\",\"content\":\"你好,介绍一下你自己\"}],\"temperature\":0.7,\"chat_template_kwargs\":{\"enable_thinking\":false}}"
|
||||
```
|
||||
|
||||
## OpenClaw 调用说明
|
||||
|
||||
- Base URL 使用 `http://<服务器IP>:8001/v1`
|
||||
- API Key 使用 `config.json` 中 `api_key`
|
||||
- 模型名使用 `config.json` 中 `models.profiles.<模型名>.served_model_name`
|
||||
- 模型名固定使用 `config.json` 中 `public_model_name`(默认 `Qwen_local_model`)
|
||||
- 思考模式按请求控制:`chat_template_kwargs.enable_thinking=false/true`
|
||||
- 若调用方未传 `chat_template_kwargs.enable_thinking`,服务端使用 `default_enable_thinking` 兜底
|
||||
- 仅当模型需要推理解析器时,再将 `config.json` 中 `reasoning_enabled` 设为 `true`
|
||||
- 若使用工具调用,`config.json` 中应配置 `tool_call_parser` 与 `enable_auto_tool_choice`
|
||||
- 服务强制离线模式,不会回退到 Hugging Face 远程下载
|
||||
- 所有路径按 Ubuntu 规范填写,本地模型建议使用 `/opt/model/<模型目录>`
|
||||
@@ -103,7 +123,10 @@ curl -X POST "http://localhost:8001/v1/chat/completions" \
|
||||
|
||||
## 常见故障排查
|
||||
|
||||
- 报错 `model type ... Transformers does not recognize this architecture` 时,说明当前模型与镜像内依赖不兼容,建议更换模型或升级镜像版本。
|
||||
- 报错 `Model architectures ['Qwen3_5MoeForConditionalGeneration'] are not supported for now` 或 `The Transformers implementation ... is not compatible with vLLM` 时,说明当前 vLLM 栈与该模型架构不兼容,需切换到兼容模型或改用其他推理后端。
|
||||
- 报错 `StrictDataclassClassValidationError` 且包含 `validate_rope` / `unsupported operand type(s) for -=: 'set' and 'list'` 时,移除 Dockerfile 中对 `transformers --upgrade --pre` 的强制升级,使用镜像内置依赖重建。
|
||||
- 报错 `moe_wna16 quantization is currently not supported in rocm` 时,将该模型的 `quantization` 改回 `gptq`。
|
||||
- 报错 `model config (gptq) does not match quantization argument (gptq_marlin)` 时,将该模型配置改为 `dtype=float16` 且 `quantization=gptq`。
|
||||
- 报错 `RPC call to sample_tokens timed out` 或出现 `GPU core dump` 时,先下调模型配置为更稳参数:`ctx=32768`、`max_num_seqs=4`、`max_tokens=2048`、`gpu_util=0.90`,并开启 `enforce_eager=true`。
|
||||
- 若模型目录存在但仍加载失败,检查挂载路径是否为 `/opt/model:/opt/model:ro`,并确认容器内可见模型文件。
|
||||
- 如果看到 `No services to build`,说明未触发重建;需要先执行 `docker compose build --no-cache` 再 `up`。
|
||||
|
||||
+45
-24
@@ -1,10 +1,9 @@
|
||||
from functools import lru_cache
|
||||
import os
|
||||
from typing import Optional
|
||||
|
||||
from pydantic import BaseModel
|
||||
|
||||
from app.model_catalog import load_catalog, resolve_model_profile, resolve_runtime_settings
|
||||
from app.model_catalog import load_app_config
|
||||
|
||||
|
||||
class Settings(BaseModel):
|
||||
@@ -17,6 +16,10 @@ class Settings(BaseModel):
|
||||
port: int = 8000
|
||||
openai_host: str = "0.0.0.0"
|
||||
openai_port: int = 8001
|
||||
vllm_openai_internal_url: str = "http://127.0.0.1:8001/v1"
|
||||
public_model_name: str = "Qwen_local_model"
|
||||
default_enable_thinking: bool = False
|
||||
reasoning_enabled: bool = False
|
||||
model_root: str = "/opt/model"
|
||||
offline_mode: bool = True
|
||||
max_model_len: int = 8192
|
||||
@@ -31,31 +34,49 @@ class Settings(BaseModel):
|
||||
enable_auto_tool_choice: bool = False
|
||||
revision: Optional[str] = None
|
||||
api_key: Optional[str] = None
|
||||
quantization: Optional[str] = None
|
||||
model_impl: Optional[str] = None
|
||||
reasoning_parser: Optional[str] = None
|
||||
kv_cache_dtype: Optional[str] = None
|
||||
enable_prefix_caching: bool = False
|
||||
max_num_batched_tokens: int = 0
|
||||
|
||||
|
||||
@lru_cache(maxsize=1)
|
||||
def get_settings() -> Settings:
|
||||
catalog = load_catalog("config.json")
|
||||
runtime = resolve_runtime_settings(catalog)
|
||||
settings = Settings(
|
||||
config = load_app_config("config.json")
|
||||
return Settings(
|
||||
config_file="config.json",
|
||||
model_key=runtime["model_key"],
|
||||
host=runtime["host"],
|
||||
port=runtime["port"],
|
||||
openai_host=runtime["openai_host"],
|
||||
openai_port=runtime["openai_port"],
|
||||
model_root=runtime["model_root"],
|
||||
offline_mode=runtime["offline_mode"],
|
||||
api_key=runtime["api_key"],
|
||||
tensor_parallel_size=runtime["tensor_parallel_size"],
|
||||
dtype=runtime["dtype"],
|
||||
revision=runtime["revision"],
|
||||
model_key=config.get("model_key"),
|
||||
selected_model=config.get("selected_model"),
|
||||
model_name=config.get("model_name", ""),
|
||||
served_model_name=config.get("served_model_name"),
|
||||
host=config.get("host", "0.0.0.0"),
|
||||
port=config.get("port", 8000),
|
||||
openai_host=config.get("openai_host", "0.0.0.0"),
|
||||
openai_port=config.get("openai_port", 8001),
|
||||
vllm_openai_internal_url=config.get("vllm_openai_internal_url", "http://127.0.0.1:8001/v1"),
|
||||
public_model_name=config.get("public_model_name", "Qwen_local_model"),
|
||||
default_enable_thinking=config.get("default_enable_thinking", False),
|
||||
reasoning_enabled=config.get("reasoning_enabled", False),
|
||||
model_root=config.get("model_root", "/opt/model"),
|
||||
offline_mode=config.get("offline_mode", True),
|
||||
max_model_len=config.get("max_model_len", 8192),
|
||||
gpu_memory_utilization=config.get("gpu_memory_utilization", 0.92),
|
||||
tensor_parallel_size=config.get("tensor_parallel_size", 2),
|
||||
max_num_seqs=config.get("max_num_seqs", 64),
|
||||
max_tokens=config.get("max_tokens", 4096),
|
||||
dtype=config.get("dtype", "bfloat16"),
|
||||
enforce_eager=config.get("enforce_eager", False),
|
||||
trust_remote_code=config.get("trust_remote_code", False),
|
||||
tool_call_parser=config.get("tool_call_parser"),
|
||||
enable_auto_tool_choice=config.get("enable_auto_tool_choice", False),
|
||||
revision=config.get("revision"),
|
||||
api_key=config.get("api_key"),
|
||||
quantization=config.get("quantization"),
|
||||
model_impl=config.get("model_impl"),
|
||||
reasoning_parser=config.get("reasoning_parser"),
|
||||
kv_cache_dtype=config.get("kv_cache_dtype"),
|
||||
enable_prefix_caching=config.get("enable_prefix_caching", False),
|
||||
max_num_batched_tokens=config.get("max_num_batched_tokens", 0),
|
||||
)
|
||||
_, updates, env_vars = resolve_model_profile(
|
||||
content=catalog,
|
||||
requested_model=settings.model_key,
|
||||
requested_tp=settings.tensor_parallel_size,
|
||||
)
|
||||
for key, value in env_vars.items():
|
||||
os.environ[key] = value
|
||||
return settings.model_copy(update=updates | runtime)
|
||||
|
||||
@@ -1,46 +0,0 @@
|
||||
import httpx
|
||||
|
||||
from app.config import Settings
|
||||
from app.schemas import GenerateRequest, GenerateResponse
|
||||
|
||||
|
||||
class InferenceEngine:
|
||||
def __init__(self, settings: Settings) -> None:
|
||||
self.settings = settings
|
||||
self.client = httpx.Client(timeout=300.0)
|
||||
|
||||
def close(self) -> None:
|
||||
self.client.close()
|
||||
|
||||
def generate(self, req: GenerateRequest) -> GenerateResponse:
|
||||
headers = {"Content-Type": "application/json"}
|
||||
if self.settings.api_key:
|
||||
headers["Authorization"] = f"Bearer {self.settings.api_key}"
|
||||
payload = {
|
||||
"model": self.settings.served_model_name or self.settings.model_name,
|
||||
"messages": [{"role": "user", "content": req.prompt}],
|
||||
"max_tokens": req.max_tokens,
|
||||
"temperature": req.temperature,
|
||||
"top_p": req.top_p,
|
||||
}
|
||||
if req.stop:
|
||||
payload["stop"] = req.stop
|
||||
response = self.client.post(
|
||||
f"{self.settings.vllm_openai_internal_url}/chat/completions",
|
||||
headers=headers,
|
||||
json=payload,
|
||||
)
|
||||
response.raise_for_status()
|
||||
body = response.json()
|
||||
completion = body["choices"][0]["message"]["content"]
|
||||
usage = body.get("usage", {})
|
||||
usage_prompt = int(usage.get("prompt_tokens", 0))
|
||||
usage_completion = int(usage.get("completion_tokens", 0))
|
||||
return GenerateResponse(
|
||||
text=completion,
|
||||
prompt=req.prompt,
|
||||
model=self.settings.served_model_name or self.settings.model_name,
|
||||
usage_prompt_tokens=usage_prompt,
|
||||
usage_completion_tokens=usage_completion,
|
||||
usage_total_tokens=usage_prompt + usage_completion,
|
||||
)
|
||||
-50
@@ -1,50 +0,0 @@
|
||||
from contextlib import asynccontextmanager
|
||||
|
||||
from fastapi import Depends, FastAPI, Header, HTTPException, status
|
||||
|
||||
from app.config import Settings, get_settings
|
||||
from app.engine import InferenceEngine
|
||||
from app.schemas import GenerateRequest, GenerateResponse, HealthResponse
|
||||
|
||||
engine: InferenceEngine | None = None
|
||||
|
||||
|
||||
def verify_api_key(
|
||||
settings: Settings = Depends(get_settings), x_api_key: str | None = Header(default=None)
|
||||
) -> None:
|
||||
if settings.api_key and x_api_key != settings.api_key:
|
||||
raise HTTPException(
|
||||
status_code=status.HTTP_401_UNAUTHORIZED,
|
||||
detail="Invalid API key",
|
||||
)
|
||||
|
||||
|
||||
@asynccontextmanager
|
||||
async def lifespan(_: FastAPI):
|
||||
global engine
|
||||
settings = get_settings()
|
||||
engine = InferenceEngine(settings)
|
||||
yield
|
||||
if engine is not None:
|
||||
engine.close()
|
||||
engine = None
|
||||
|
||||
|
||||
app = FastAPI(title="ROCm vLLM Inference API", version="1.0.0", lifespan=lifespan)
|
||||
|
||||
|
||||
@app.get("/health", response_model=HealthResponse)
|
||||
def health(settings: Settings = Depends(get_settings)) -> HealthResponse:
|
||||
return HealthResponse(status="ok", model=settings.served_model_name or settings.model_name)
|
||||
|
||||
|
||||
@app.post("/v1/generate", response_model=GenerateResponse, dependencies=[Depends(verify_api_key)])
|
||||
def generate(req: GenerateRequest, settings: Settings = Depends(get_settings)) -> GenerateResponse:
|
||||
if engine is None:
|
||||
raise HTTPException(status_code=status.HTTP_503_SERVICE_UNAVAILABLE, detail="Engine not ready")
|
||||
if req.max_tokens > settings.max_tokens:
|
||||
raise HTTPException(
|
||||
status_code=status.HTTP_422_UNPROCESSABLE_ENTITY,
|
||||
detail=f"max_tokens must be <= {settings.max_tokens}",
|
||||
)
|
||||
return engine.generate(req)
|
||||
+71
-1
@@ -1,4 +1,5 @@
|
||||
import json
|
||||
import os
|
||||
from pathlib import Path
|
||||
from typing import Any
|
||||
|
||||
@@ -63,12 +64,32 @@ def resolve_runtime_settings(content: dict[str, Any]) -> dict[str, Any]:
|
||||
api_service = dict(services.get("api", {}))
|
||||
openai_service = dict(services.get("openai", {}))
|
||||
models = dict(content.get("models", {}))
|
||||
openai_port = _to_int(openai_service.get("port"), 8001)
|
||||
internal_url = _to_str(content.get("vllm_openai_internal_url"))
|
||||
if not internal_url:
|
||||
internal_url = f"http://127.0.0.1:{openai_port}/v1"
|
||||
enable_thinking_env = os.getenv("VLLM_ENABLE_THINKING")
|
||||
if enable_thinking_env is not None:
|
||||
default_enable_thinking = _to_bool(enable_thinking_env, False)
|
||||
else:
|
||||
default_enable_thinking = _to_bool(content.get("default_enable_thinking"), False)
|
||||
|
||||
reasoning_enabled_env = os.getenv("VLLM_REASONING_ENABLED")
|
||||
if reasoning_enabled_env is not None:
|
||||
reasoning_enabled = _to_bool(reasoning_enabled_env, False)
|
||||
else:
|
||||
reasoning_enabled = _to_bool(content.get("reasoning_enabled"), False)
|
||||
|
||||
return {
|
||||
"host": str(api_service.get("host", "0.0.0.0")),
|
||||
"port": _to_int(api_service.get("port"), 8000),
|
||||
"openai_host": str(openai_service.get("host", "0.0.0.0")),
|
||||
"openai_port": _to_int(openai_service.get("port"), 8001),
|
||||
"openai_port": openai_port,
|
||||
"vllm_openai_internal_url": internal_url.rstrip("/"),
|
||||
"public_model_name": _to_str(content.get("public_model_name"), "Qwen_local_model"),
|
||||
"default_enable_thinking": default_enable_thinking,
|
||||
"api_key": str(content.get("api_key", "")).strip() or None,
|
||||
"reasoning_enabled": reasoning_enabled,
|
||||
"tensor_parallel_size": _to_int(content.get("tensor_parallel_size"), 2),
|
||||
"dtype": str(content.get("dtype", "bfloat16")),
|
||||
"revision": str(content.get("revision", "")).strip() or None,
|
||||
@@ -96,12 +117,27 @@ def resolve_model_profile(
|
||||
resolved_tp = requested_tp
|
||||
if valid_tp and resolved_tp not in valid_tp:
|
||||
resolved_tp = valid_tp[0]
|
||||
speculative = dict(profile.get("speculative", {}))
|
||||
speculative_method = _to_str(speculative.get("method"))
|
||||
speculative_model_path = _to_str(speculative.get("model"))
|
||||
num_speculative_tokens = _to_int(speculative.get("num_speculative_tokens"), 0)
|
||||
speculative_draft_tp = _to_int(speculative.get("draft_tensor_parallel_size"), 0)
|
||||
if speculative_method and speculative_model_path:
|
||||
resolved_speculative_model = speculative_model_path
|
||||
if not speculative_model_path.startswith("/"):
|
||||
if not model_root:
|
||||
raise ValueError("config.json model_root cannot be empty when speculative model path is relative")
|
||||
resolved_speculative_model = _join_posix(model_root, speculative_model_path)
|
||||
else:
|
||||
resolved_speculative_model = ""
|
||||
updates = {
|
||||
"selected_model": model_key,
|
||||
"model_name": _resolve_profile_model_path(profile, model_root, model_key),
|
||||
"served_model_name": profile.get("served_model_name", model_key),
|
||||
"dtype": _to_str(profile.get("dtype")),
|
||||
"quantization": _to_str(profile.get("quantization")),
|
||||
"model_impl": _to_str(profile.get("model_impl")),
|
||||
"reasoning_parser": _to_str(os.getenv("VLLM_REASONING_PARSER") or profile.get("reasoning_parser")),
|
||||
"max_model_len": _to_int(profile.get("ctx"), 8192),
|
||||
"max_num_seqs": _to_int(profile.get("max_num_seqs"), 64),
|
||||
"max_tokens": _to_int(profile.get("max_tokens"), 4096),
|
||||
@@ -111,8 +147,42 @@ def resolve_model_profile(
|
||||
"tensor_parallel_size": resolved_tp,
|
||||
"tool_call_parser": profile.get("tool_call_parser"),
|
||||
"enable_auto_tool_choice": _to_bool(profile.get("enable_auto_tool_choice"), False),
|
||||
"kv_cache_dtype": _to_str(profile.get("kv_cache_dtype")),
|
||||
"enable_prefix_caching": _to_bool(profile.get("enable_prefix_caching"), False),
|
||||
"max_num_batched_tokens": _to_int(profile.get("max_num_batched_tokens"), 0),
|
||||
"language_model_only": _to_bool(profile.get("language_model_only"), False),
|
||||
"speculative_method": speculative_method,
|
||||
"speculative_model": resolved_speculative_model,
|
||||
"num_speculative_tokens": num_speculative_tokens,
|
||||
"speculative_draft_tp": speculative_draft_tp,
|
||||
}
|
||||
env_vars = {str(k): str(v) for k, v in dict(profile.get("env", {})).items()}
|
||||
env_vars["HF_HUB_OFFLINE"] = "1"
|
||||
env_vars["TRANSFORMERS_OFFLINE"] = "1"
|
||||
if "VLLM_RPC_TIMEOUT" not in env_vars:
|
||||
env_vars["VLLM_RPC_TIMEOUT"] = "300"
|
||||
if "VLLM_WORKER_MULTIPROC_METHOD" not in env_vars:
|
||||
env_vars["VLLM_WORKER_MULTIPROC_METHOD"] = "spawn"
|
||||
return model_key, updates, env_vars
|
||||
|
||||
|
||||
def load_app_config(config_file: str = "config.json") -> dict[str, Any]:
|
||||
"""
|
||||
统一的应用配置加载函数,封装完整的配置加载流程。
|
||||
|
||||
Args:
|
||||
config_file: 配置文件路径,默认为 "config.json"
|
||||
|
||||
Returns:
|
||||
包含合并后配置的字典,包括 runtime settings 和 model profile updates
|
||||
"""
|
||||
catalog = load_catalog(config_file)
|
||||
runtime = resolve_runtime_settings(catalog)
|
||||
_, updates, env_vars = resolve_model_profile(
|
||||
content=catalog,
|
||||
requested_model=runtime["model_key"],
|
||||
requested_tp=runtime["tensor_parallel_size"],
|
||||
)
|
||||
for key, value in env_vars.items():
|
||||
os.environ[key] = value
|
||||
return {**runtime, **updates}
|
||||
|
||||
@@ -1,26 +0,0 @@
|
||||
from typing import List, Optional
|
||||
|
||||
from pydantic import BaseModel, Field
|
||||
|
||||
|
||||
class GenerateRequest(BaseModel):
|
||||
prompt: str
|
||||
max_tokens: int = Field(default=256, ge=1, le=4096)
|
||||
temperature: float = Field(default=0.7, ge=0.0, le=2.0)
|
||||
top_p: float = Field(default=0.95, gt=0.0, le=1.0)
|
||||
repetition_penalty: float = Field(default=1.0, ge=0.5, le=2.0)
|
||||
stop: Optional[List[str]] = None
|
||||
|
||||
|
||||
class GenerateResponse(BaseModel):
|
||||
text: str
|
||||
prompt: str
|
||||
model: str
|
||||
usage_prompt_tokens: int
|
||||
usage_completion_tokens: int
|
||||
usage_total_tokens: int
|
||||
|
||||
|
||||
class HealthResponse(BaseModel):
|
||||
status: str
|
||||
model: str
|
||||
@@ -1,23 +0,0 @@
|
||||
import subprocess
|
||||
import sys
|
||||
|
||||
from app.config import get_settings
|
||||
|
||||
|
||||
def main() -> None:
|
||||
settings = get_settings()
|
||||
command = [
|
||||
sys.executable,
|
||||
"-m",
|
||||
"uvicorn",
|
||||
"app.main:app",
|
||||
"--host",
|
||||
settings.host,
|
||||
"--port",
|
||||
str(settings.port),
|
||||
]
|
||||
raise SystemExit(subprocess.call(command))
|
||||
|
||||
|
||||
if __name__ == "__main__":
|
||||
main()
|
||||
+56
-29
@@ -1,27 +1,23 @@
|
||||
import os
|
||||
import json
|
||||
import subprocess
|
||||
import sys
|
||||
|
||||
from app.model_catalog import load_catalog, resolve_model_profile, resolve_runtime_settings
|
||||
from app.model_catalog import load_app_config
|
||||
|
||||
|
||||
def build_command() -> list[str]:
|
||||
config_file = "config.json"
|
||||
catalog = load_catalog(config_file)
|
||||
runtime = resolve_runtime_settings(catalog)
|
||||
_, updates, env_vars = resolve_model_profile(
|
||||
content=catalog,
|
||||
requested_model=runtime["model_key"],
|
||||
requested_tp=runtime["tensor_parallel_size"],
|
||||
)
|
||||
for key, value in env_vars.items():
|
||||
os.environ[key] = value
|
||||
host = str(runtime["openai_host"])
|
||||
port = str(runtime["openai_port"])
|
||||
api_key = runtime["api_key"] or ""
|
||||
dtype = str(updates["dtype"] or runtime["dtype"])
|
||||
quantization = str(updates["quantization"] or "").strip()
|
||||
revision = runtime["revision"] or ""
|
||||
config = load_app_config("config.json")
|
||||
host = str(config["openai_host"])
|
||||
port = str(config["openai_port"])
|
||||
public_model_name = str(config["public_model_name"]).strip()
|
||||
default_enable_thinking = bool(config["default_enable_thinking"])
|
||||
reasoning_enabled = bool(config["reasoning_enabled"])
|
||||
api_key = config.get("api_key") or ""
|
||||
dtype = str(config.get("dtype", "bfloat16"))
|
||||
quantization = str(config.get("quantization", "")).strip()
|
||||
model_impl = str(config.get("model_impl", "")).strip()
|
||||
reasoning_parser = str(config.get("reasoning_parser", "")).strip()
|
||||
revision = config.get("revision", "") or ""
|
||||
cmd = [
|
||||
sys.executable,
|
||||
"-m",
|
||||
@@ -31,34 +27,65 @@ def build_command() -> list[str]:
|
||||
"--port",
|
||||
port,
|
||||
"--model",
|
||||
str(updates["model_name"]),
|
||||
str(config["model_name"]),
|
||||
"--served-model-name",
|
||||
str(updates["served_model_name"]),
|
||||
public_model_name or str(config["served_model_name"]),
|
||||
"--tensor-parallel-size",
|
||||
str(updates["tensor_parallel_size"]),
|
||||
str(config["tensor_parallel_size"]),
|
||||
"--max-model-len",
|
||||
str(updates["max_model_len"]),
|
||||
str(config["max_model_len"]),
|
||||
"--gpu-memory-utilization",
|
||||
str(updates["gpu_memory_utilization"]),
|
||||
str(config["gpu_memory_utilization"]),
|
||||
"--max-num-seqs",
|
||||
str(updates["max_num_seqs"]),
|
||||
str(config["max_num_seqs"]),
|
||||
"--dtype",
|
||||
dtype,
|
||||
]
|
||||
if updates["trust_remote_code"]:
|
||||
if config["trust_remote_code"]:
|
||||
cmd.append("--trust-remote-code")
|
||||
if updates["enforce_eager"]:
|
||||
if config["enforce_eager"]:
|
||||
cmd.append("--enforce-eager")
|
||||
if updates["enable_auto_tool_choice"]:
|
||||
if config["enable_auto_tool_choice"]:
|
||||
cmd.append("--enable-auto-tool-choice")
|
||||
if updates["tool_call_parser"]:
|
||||
cmd.extend(["--tool-call-parser", str(updates["tool_call_parser"])])
|
||||
if config["tool_call_parser"]:
|
||||
cmd.extend(["--tool-call-parser", str(config["tool_call_parser"])])
|
||||
cmd.extend(
|
||||
[
|
||||
"--default-chat-template-kwargs",
|
||||
json.dumps({"enable_thinking": default_enable_thinking}),
|
||||
]
|
||||
)
|
||||
if reasoning_enabled and reasoning_parser:
|
||||
cmd.extend(["--reasoning-parser", reasoning_parser])
|
||||
if quantization:
|
||||
cmd.extend(["--quantization", quantization])
|
||||
if model_impl:
|
||||
cmd.extend(["--model-impl", model_impl])
|
||||
if revision:
|
||||
cmd.extend(["--revision", revision])
|
||||
if api_key:
|
||||
cmd.extend(["--api-key", api_key])
|
||||
if config.get("kv_cache_dtype"):
|
||||
cmd.extend(["--kv-cache-dtype", str(config["kv_cache_dtype"])])
|
||||
if config.get("enable_prefix_caching"):
|
||||
cmd.append("--enable-prefix-caching")
|
||||
if config.get("max_num_batched_tokens", 0) > 0:
|
||||
cmd.extend(["--max-num-batched-tokens", str(config["max_num_batched_tokens"])])
|
||||
if config.get("language_model_only"):
|
||||
cmd.append("--language-model-only")
|
||||
speculative_method = str(config.get("speculative_method") or "").strip()
|
||||
speculative_model = str(config.get("speculative_model") or "").strip()
|
||||
num_speculative_tokens = int(config.get("num_speculative_tokens") or 0)
|
||||
speculative_draft_tp = int(config.get("speculative_draft_tp") or 0)
|
||||
if speculative_method and speculative_model and num_speculative_tokens > 0:
|
||||
spec_config: dict = {
|
||||
"method": speculative_method,
|
||||
"model": speculative_model,
|
||||
"num_speculative_tokens": num_speculative_tokens,
|
||||
}
|
||||
if speculative_draft_tp > 0:
|
||||
spec_config["draft_tensor_parallel_size"] = speculative_draft_tp
|
||||
cmd.extend(["--speculative-config", json.dumps(spec_config)])
|
||||
return cmd
|
||||
|
||||
|
||||
|
||||
@@ -0,0 +1,99 @@
|
||||
#!/usr/bin/env python3
|
||||
"""
|
||||
交互式对话脚本,用于与vLLM模型进行对话并计算token生成速度
|
||||
|
||||
使用方法:
|
||||
1. 进入Docker容器:docker exec -it rocm-vllm-openai bash
|
||||
2. 运行:python chat_with_speed.py
|
||||
3. 输入提示词与模型对话
|
||||
4. 输入 'exit' 退出
|
||||
"""
|
||||
|
||||
import json
|
||||
import time
|
||||
import httpx
|
||||
|
||||
# 模型服务地址
|
||||
API_URL = "http://localhost:8001/v1/chat/completions"
|
||||
# API密钥
|
||||
API_KEY = "sk-szcjw"
|
||||
# 模型名称
|
||||
MODEL_NAME = "Qwen_local_model"
|
||||
|
||||
def chat_with_model():
|
||||
"""交互式对话函数"""
|
||||
print("=== vLLM 交互式对话工具 ===")
|
||||
print("输入提示词与模型对话,输入 'exit' 退出")
|
||||
print("=" * 50)
|
||||
|
||||
# 对话历史
|
||||
messages = []
|
||||
|
||||
while True:
|
||||
# 获取用户输入
|
||||
user_input = input("用户: ").strip()
|
||||
|
||||
if user_input.lower() == "exit":
|
||||
print("退出对话...")
|
||||
break
|
||||
|
||||
if not user_input:
|
||||
continue
|
||||
|
||||
# 添加用户消息到对话历史
|
||||
messages.append({"role": "user", "content": user_input})
|
||||
|
||||
# 准备请求数据
|
||||
payload = {
|
||||
"model": MODEL_NAME,
|
||||
"messages": messages,
|
||||
"max_tokens": 1000,
|
||||
"temperature": 0.7,
|
||||
"top_p": 0.8,
|
||||
"top_k": 20
|
||||
}
|
||||
|
||||
headers = {
|
||||
"Content-Type": "application/json",
|
||||
"Authorization": f"Bearer {API_KEY}"
|
||||
}
|
||||
|
||||
print("模型: ", end="", flush=True)
|
||||
|
||||
# 记录开始时间
|
||||
start_time = time.time()
|
||||
|
||||
try:
|
||||
# 发送请求
|
||||
response = httpx.post(API_URL, json=payload, headers=headers, timeout=300.0)
|
||||
response.raise_for_status()
|
||||
|
||||
# 解析响应
|
||||
result = response.json()
|
||||
|
||||
# 获取模型回复
|
||||
assistant_message = result["choices"][0]["message"]["content"]
|
||||
print(assistant_message)
|
||||
|
||||
# 添加模型回复到对话历史
|
||||
messages.append({"role": "assistant", "content": assistant_message})
|
||||
|
||||
# 计算token速度
|
||||
usage = result.get("usage", {})
|
||||
completion_tokens = usage.get("completion_tokens", 0)
|
||||
end_time = time.time()
|
||||
elapsed_time = end_time - start_time
|
||||
|
||||
if completion_tokens > 0 and elapsed_time > 0:
|
||||
tokens_per_second = completion_tokens / elapsed_time
|
||||
print(f"\n[速度统计] 生成 {completion_tokens} tokens,用时 {elapsed_time:.2f} 秒,速度: {tokens_per_second:.2f} tokens/s")
|
||||
else:
|
||||
print("\n[速度统计] 无法计算速度")
|
||||
|
||||
except Exception as e:
|
||||
print(f"\n错误: {e}")
|
||||
|
||||
print("=" * 50)
|
||||
|
||||
if __name__ == "__main__":
|
||||
chat_with_model()
|
||||
+70
-66
@@ -9,6 +9,9 @@
|
||||
"port": 8001
|
||||
}
|
||||
},
|
||||
"public_model_name": "Qwen_local_model",
|
||||
"default_enable_thinking": false,
|
||||
"reasoning_enabled": false,
|
||||
"api_key": "sk-szcjw",
|
||||
"tensor_parallel_size": 2,
|
||||
"dtype": "bfloat16",
|
||||
@@ -17,79 +20,80 @@
|
||||
"revision": "",
|
||||
"models": {
|
||||
"default": "Qwen3.5-35B-A3B-GPTQ-Int4",
|
||||
"selected": "Qwen3.5-35B-A3B-GPTQ-Int4",
|
||||
"selected": "Qwen3.6-27B-FP8",
|
||||
"profiles": {
|
||||
"Qwen3-Next-80B-A3B-Instruct-AWQ-4bit": {
|
||||
"local_path": "Qwen3-Next-80B-A3B-Instruct-AWQ-4bit",
|
||||
"ctx": "24576",
|
||||
"trust_remote": true,
|
||||
"valid_tp": [2],
|
||||
"max_num_seqs": "32",
|
||||
"max_tokens": "16384",
|
||||
"gpu_util": "0.98",
|
||||
"enforce_eager": false,
|
||||
"env": {
|
||||
"VLLM_USE_TRITON_AWQ": "1"
|
||||
},
|
||||
"tool_call_parser": "qwen3_xml",
|
||||
"enable_auto_tool_choice": true,
|
||||
"served_model_name": "Qwen3-Next-80B-A3B-Instruct-AWQ-4bit",
|
||||
"hf_model_id": "cpatonn/Qwen3-Next-80B-A3B-Instruct-AWQ-4bit"
|
||||
},
|
||||
"GLM-4.7-Flash-AWQ": {
|
||||
"local_path": "GLM-4.7-Flash-AWQ",
|
||||
"ctx": "32768",
|
||||
"trust_remote": true,
|
||||
"valid_tp": [1, 2],
|
||||
"max_num_seqs": "64",
|
||||
"max_tokens": "32768",
|
||||
"gpu_util": "0.98",
|
||||
"tool_call_parser": "qwen3_xml",
|
||||
"enable_auto_tool_choice": true,
|
||||
"served_model_name": "GLM-4.7-Flash-AWQ",
|
||||
"hf_model_id": "THUDM/GLM-4.7-Flash-AWQ"
|
||||
},
|
||||
"Qwen3.5-27B-FP8": {
|
||||
"local_path": "Qwen3.5-27B-FP8",
|
||||
"ctx": "65536",
|
||||
"trust_remote": true,
|
||||
"valid_tp": [1, 2],
|
||||
"max_num_seqs": "64",
|
||||
"max_tokens": "32768",
|
||||
"gpu_util": "0.98",
|
||||
"tool_call_parser": "qwen3_xml",
|
||||
"enable_auto_tool_choice": true,
|
||||
"served_model_name": "Qwen3.5-27B-FP8",
|
||||
"hf_model_id": "RedHatAI/Qwen3.5-27B-FP8-dynamic"
|
||||
},
|
||||
"Qwen3.5-35B-A3B-GPTQ-Int4": {
|
||||
"local_path": "Qwen3.5-35B-A3B-GPTQ-Int4",
|
||||
"Qwen3.6-35B-A3B-FP8": {
|
||||
"local_path": "Qwen3.6-35B-A3B-FP8",
|
||||
"dtype": "float16",
|
||||
"quantization": "gptq",
|
||||
"ctx": "65536",
|
||||
"quantization": "fp8",
|
||||
"ctx": "262144",
|
||||
"max_tokens": "65536",
|
||||
"max_num_batched_tokens": 32768,
|
||||
"trust_remote": true,
|
||||
"valid_tp": [1, 2],
|
||||
"max_num_seqs": "64",
|
||||
"max_tokens": "32768",
|
||||
"gpu_util": "0.98",
|
||||
"tool_call_parser": "qwen3_xml",
|
||||
"enforce_eager": false,
|
||||
"valid_tp": [
|
||||
2
|
||||
],
|
||||
"max_num_seqs": "256",
|
||||
"gpu_util": "0.92",
|
||||
"tool_call_parser": "qwen3_coder",
|
||||
"reasoning_parser": "qwen3",
|
||||
"enable_auto_tool_choice": true,
|
||||
"served_model_name": "Qwen3.5-35B-A3B-GPTQ-Int4",
|
||||
"hf_model_id": "Qwen/Qwen3.5-35B-A3B-GPTQ-Int4"
|
||||
"language_model_only": true,
|
||||
"served_model_name": "Qwen3.6-35B-A3B-FP8",
|
||||
"hf_model_id": "Qwen/Qwen3.6-35B-A3B-FP8"
|
||||
},
|
||||
"Qwen3.5-35B-A3B-FP8": {
|
||||
"local_path": "Qwen3.5-35B-A3B-FP8",
|
||||
"dtype": "bfloat16",
|
||||
"ctx": "65536",
|
||||
"trust_remote": true,
|
||||
"valid_tp": [1, 2],
|
||||
"max_num_seqs": "64",
|
||||
"Qwen3.6-27B": {
|
||||
"local_path": "Qwen3.6-27B",
|
||||
"ctx": "131072",
|
||||
"max_tokens": "32768",
|
||||
"gpu_util": "0.98",
|
||||
"tool_call_parser": "qwen3_xml",
|
||||
"max_num_batched_tokens": 16384,
|
||||
"trust_remote": true,
|
||||
"enforce_eager": false,
|
||||
"valid_tp": [
|
||||
1,
|
||||
2
|
||||
],
|
||||
"max_num_seqs": 32,
|
||||
"gpu_util": "0.96",
|
||||
"kv_cache_dtype": "fp8",
|
||||
"enable_prefix_caching": true,
|
||||
"tool_call_parser": "qwen3_coder",
|
||||
"reasoning_parser": "qwen3",
|
||||
"enable_auto_tool_choice": true,
|
||||
"served_model_name": "Qwen3.5-35B-A3B-FP8",
|
||||
"hf_model_id": "Qwen/Qwen3.5-35B-A3B-FP8"
|
||||
"language_model_only": true,
|
||||
"served_model_name": "Qwen3.6-27B",
|
||||
"hf_model_id": "Qwen/Qwen3.6-27B"
|
||||
},
|
||||
"Qwen3.6-27B-FP8": {
|
||||
"local_path": "Qwen3.6-27B-FP8",
|
||||
"dtype": "auto",
|
||||
"quantization": "fp8",
|
||||
"ctx": "131072",
|
||||
"max_tokens": "32768",
|
||||
"max_num_batched_tokens": 16384,
|
||||
"trust_remote": true,
|
||||
"enforce_eager": false,
|
||||
"valid_tp": [
|
||||
2
|
||||
],
|
||||
"max_num_seqs": 32,
|
||||
"gpu_util": "0.85",
|
||||
"tool_call_parser": "qwen3_coder",
|
||||
"reasoning_parser": "qwen3",
|
||||
"enable_auto_tool_choice": true,
|
||||
"language_model_only": true,
|
||||
"served_model_name": "Qwen3.6-27B-FP8",
|
||||
"hf_model_id": "Qwen/Qwen3.6-27B-FP8",
|
||||
"env": {
|
||||
"VLLM_RPC_TIMEOUT": "300"
|
||||
},
|
||||
"speculative": {
|
||||
"method": "dflash",
|
||||
"model": "Qwen3.6-27B-DFlash",
|
||||
"num_speculative_tokens": 4,
|
||||
"draft_tensor_parallel_size": 1
|
||||
}
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
+16
-3
@@ -3,8 +3,9 @@ services:
|
||||
build:
|
||||
context: .
|
||||
dockerfile: Dockerfile
|
||||
image: rocm-vllm-inference:latest
|
||||
container_name: rocm-vllm-openai
|
||||
pull: false
|
||||
image: vllm-openai-rocm:nightly
|
||||
container_name: vllm-openai-rocm
|
||||
entrypoint: ["python"]
|
||||
command: ["-m", "app.start_openai"]
|
||||
ports:
|
||||
@@ -12,13 +13,25 @@ services:
|
||||
volumes:
|
||||
- /opt/model:/opt/model:ro
|
||||
- ./config.json:/workspace/config.json:ro
|
||||
environment:
|
||||
HF_HOME: /opt/model
|
||||
# 严格离线模式,禁止任何网络下载
|
||||
HF_HUB_OFFLINE: "1"
|
||||
TRANSFORMERS_OFFLINE: "1"
|
||||
HF_DATASETS_OFFLINE: "1"
|
||||
VLLM_ROCM_USE_AITER: "0" # 打开总开关
|
||||
# 明确禁用不稳定的MoE后端,让其他AITER优化(如MHA)生效
|
||||
VLLM_ROCM_USE_AITER_MOE: "0"
|
||||
VLLM_ROCM_MOE_BACKEND: "TRITON" # 强制MoE使用Triton
|
||||
TORCHINDUCTOR_FX_GRAPH_CACHE: "0" # 禁用 torch.compile 缓存
|
||||
VLLM_DISABLE_COMPILE_CACHE: "1"
|
||||
devices:
|
||||
- /dev/kfd
|
||||
- /dev/dri
|
||||
group_add:
|
||||
- video
|
||||
ipc: host
|
||||
shm_size: 16g
|
||||
shm_size: 32g
|
||||
cap_add:
|
||||
- SYS_PTRACE
|
||||
security_opt:
|
||||
|
||||
Binary file not shown.
Binary file not shown.
Binary file not shown.
Binary file not shown.
Binary file not shown.
Binary file not shown.
Binary file not shown.
Reference in New Issue
Block a user