> ## Documentation Index
> Fetch the complete documentation index at: https://docs.hhapi.xyz/llms.txt
> Use this file to discover all available pages before exploring further.

# POST /v1/chat/completions — HHAPI 对话生成接口：请求参数与响应字段完整说明

> POST /v1/chat/completions 根据 messages 列表生成 AI 回复，支持 model、stream、max_tokens、temperature、top_p 等参数控制。调用会消耗账户余额，OpenAI 参数兼容范围待确认，请以本文档所列参数为准。

`POST /v1/chat/completions` 是 HHAPI 的核心接口，用于基于输入消息列表生成 AI 模型的回复。支持单轮问答与多轮对话，可通过系统提示词（`system` 角色消息）设定模型的行为风格，并提供多种参数用于控制生成质量与长度。接口设计参考 OpenAI Chat Completions 规范，**具体参数兼容范围待确认**，以本文档所列参数为准。

<Warning>
  调用此接口会**消耗账户余额**。请确保账户余额充足，并注意控制 `max_tokens` 等参数以避免非预期的高额费用。
</Warning>

<Note>
  **接口信息**

  * **方法**：`POST`
  * **路径**：`/v1/chat/completions`
  * **鉴权**：必需（Bearer Token）
</Note>

## 请求参数

<ParamField body="model" type="string" required>
  要使用的模型 ID。可通过 [GET /v1/models](/api-reference/models) 接口获取当前可用的模型列表。
</ParamField>

<ParamField body="messages" type="array" required>
  对话消息数组，按时间顺序排列，构成完整的对话上下文。每个消息对象需包含 `role` 和 `content` 字段。

  <Expandable title="messages 对象字段">
    <ParamField body="role" type="string" required>
      消息角色，可选值：

      * `"system"` — 系统提示词，用于设定模型行为和背景
      * `"user"` — 用户输入的消息
      * `"assistant"` — 模型此前的回复（用于多轮对话上下文）
    </ParamField>

    <ParamField body="content" type="string" required>
      消息的文本内容。
    </ParamField>
  </Expandable>
</ParamField>

<ParamField body="stream" type="boolean" default="false">
  是否启用流式响应（Server-Sent Events）。设为 `true` 时，模型将逐块返回生成内容而非等待全部生成完毕。详见 [流式响应说明](/api-reference/streaming)。
</ParamField>

<ParamField body="max_tokens" type="integer">
  本次请求允许生成的最大 Token 数量。超出此限制后生成将被截断，`finish_reason` 返回 `"length"`。
</ParamField>

<ParamField body="temperature" type="number">
  采样温度，范围 `0` 到 `2`。值越高，输出越随机多样；值越低，输出越确定集中。通常建议使用默认值或在 `0.7`–`1.0` 之间调整。不建议同时修改 `temperature` 和 `top_p`。
</ParamField>

<ParamField body="top_p" type="number">
  核采样（nucleus sampling）参数，范围 `0` 到 `1`。模型只从累积概率达到 `top_p` 的最小 token 集合中采样。不建议同时修改 `temperature` 和 `top_p`。
</ParamField>

<ParamField body="n" type="integer" default="1">
  为每个输入消息生成的回复条数。**注意：此参数是否受支持待确认，建议默认使用 `1`。**
</ParamField>

<ParamField body="stop" type="string | array">
  停止序列。当模型生成内容中出现指定字符串时，停止继续生成。可传入单个字符串或字符串数组（最多 4 个）。**注意：此参数是否受支持待确认。**
</ParamField>

## 请求示例

以下示例从环境变量读取 API Key 和 Base URL：

```bash theme={null}
curl "$API_BASE_URL/v1/chat/completions" \
  -H "Content-Type: application/json" \
  -H "Authorization: Bearer $HHAPI_API_KEY" \
  -d '{
    "model": "<MODEL_ID>",
    "messages": [
      {"role": "system", "content": "你是一个助手。"},
      {"role": "user", "content": "你好"}
    ],
    "max_tokens": 512
  }'
```

## 响应字段

<ResponseField name="id" type="string">
  本次请求的唯一 ID，格式通常为 `chatcmpl-` 开头的字符串，可用于日志追踪。
</ResponseField>

<ResponseField name="object" type="string">
  固定值 `"chat.completion"`。
</ResponseField>

<ResponseField name="created" type="integer">
  响应创建时的 Unix 时间戳（秒级）。
</ResponseField>

<ResponseField name="model" type="string">
  本次请求实际使用的模型 ID。
</ResponseField>

<ResponseField name="choices" type="array">
  模型回复列表。默认情况下（`n=1`）仅包含一个元素。
</ResponseField>

<ResponseField name="choices[].message.role" type="string">
  回复消息的角色，固定为 `"assistant"`。
</ResponseField>

<ResponseField name="choices[].message.content" type="string">
  模型生成的回复文本内容。
</ResponseField>

<ResponseField name="choices[].finish_reason" type="string">
  生成结束的原因：

  * `"stop"` — 自然结束或触发停止序列
  * `"length"` — 达到 `max_tokens` 限制
  * `null` — 流式模式中间块
</ResponseField>

<ResponseField name="usage.prompt_tokens" type="integer">
  输入消息（prompt）消耗的 Token 数量。
</ResponseField>

<ResponseField name="usage.completion_tokens" type="integer">
  模型生成回复消耗的 Token 数量。
</ResponseField>

<ResponseField name="usage.total_tokens" type="integer">
  本次请求消耗的总 Token 数量（prompt\_tokens + completion\_tokens）。
</ResponseField>

## 响应示例

```json theme={null}
{
  "id": "chatcmpl-xxxxxxxxxxxx",
  "object": "chat.completion",
  "created": 1720000000,
  "model": "<MODEL_ID>",
  "choices": [
    {
      "index": 0,
      "message": {
        "role": "assistant",
        "content": "你好！有什么我可以帮助你的吗？"
      },
      "finish_reason": "stop"
    }
  ],
  "usage": {
    "prompt_tokens": 20,
    "completion_tokens": 15,
    "total_tokens": 35
  }
}
```

## 错误说明

| 状态码 | 常见原因 |
| - | - |
| `400` | 请求体格式错误，或缺少必需字段（如 `model`、`messages`） |
| `401` | API Key 缺失或无效 |
| `429` | 超出请求速率限制 |
| `500` | 服务端内部错误 |

完整的错误码和错误处理说明请参阅 [错误响应说明](/api-reference/errors)。
