Documentazione

Tutto quello che serve per connettere la tua applicazione all'API ReQurv Hive, compatibile con OpenAI.

Dettagli di connessione

Usa questi valori in qualsiasi client compatibile con OpenAI.

URL base

https://hive.requrv.ai/api/v1

Intestazione di autenticazione

Authorization: Bearer requrv_sk_…

ReQurv Proxy — OpenAI-Compatible API

ReQurv exposes your organization's LLM models through a drop-in, OpenAI-compatible API. If your tool, framework, or custom code can talk to the OpenAI API, it can talk to ReQurv — you only need to change the base URL and the API key.

The proxy supports chat completions, legacy text completions, embeddings, audio transcription, and speech generation, with both synchronous and streaming responses.


1. Prerequisites

Before you can make your first request you need an API key.

  1. Open the API Keys page from the sidebar.
  2. Click Crea chiave (Create key) and give the key a name.
  3. Copy the secret immediately — it is shown only once.

API keys always start with requrv_sk_ and are sent in the Authorization header:

Authorization: Bearer requrv_sk_xxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxx
Treat your API key like a password. Anyone who has it can spend your organization's quota. If a key is compromised, revoke it from the API Keys page and create a new one.

2. Connection details

SettingValue
Base URLhttps://your-requrv-instance.com/api/v1
AuthenticationAuthorization: Bearer <YOUR_API_KEY>
ModelsListed by GET /v1/models (see below)

Replace https://your-requrv-instance.com with the address of your ReQurv instance (the same URL you use to open the dashboard). The OpenAI-compatible endpoints live under the /api/v1 path.

The model name you send in requests must match one of the models configured for your organization. To see the exact names available to you:

curl https://your-requrv-instance.com/api/v1/models \
  -H "Authorization: Bearer requrv_sk_YOUR_API_KEY"
The proxy is a gateway: it forwards requests to the model servers configured for your organization and records usage. It does not host the models itself.

3. Quick start

The simplest possible chat completion with curl:

curl
curl https://your-requrv-instance.com/api/v1/chat/completions \
  -H "Content-Type: application/json" \
  -H "Authorization: Bearer requrv_sk_YOUR_API_KEY" \
  -d '{
    "model": "your-model-name",
    "messages": [
      { "role": "system", "content": "You are a helpful assistant." },
      { "role": "user", "content": "Explain the proxy in one sentence." }
    ]
  }'

A successful response looks like this:

Response
{
  "id": "chatcmpl-...",
  "object": "chat.completion",
  "created": 1721234567,
  "model": "your-model-name",
  "choices": [
    {
      "index": 0,
      "message": {
        "role": "assistant",
        "content": "The ReQurv proxy is an OpenAI-compatible gateway that lets you call your organization's LLM models from any tool or framework."
      },
      "finish_reason": "stop"
    }
  ],
  "usage": {
    "prompt_tokens": 24,
    "completion_tokens": 22,
    "total_tokens": 46
  }
}

4. Code examples

All examples use https://your-requrv-instance.com/api/v1 as the base URL and requrv_sk_YOUR_API_KEY as the key. Replace both with your own values.

4.1 Python — openai SDK

python
from openai import OpenAI

client = OpenAI(
    base_url="https://your-requrv-instance.com/api/v1",
    api_key="requrv_sk_YOUR_API_KEY",
)

response = client.chat.completions.create(
    model="your-model-name",
    messages=[
        {"role": "system", "content": "You are a concise assistant."},
        {"role": "user", "content": "What is the capital of France?"},
    ],
)

print(response.choices[0].message.content)

4.2 Python — streaming with the openai SDK

python
from openai import OpenAI

client = OpenAI(
    base_url="https://your-requrv-instance.com/api/v1",
    api_key="requrv_sk_YOUR_API_KEY",
)

stream = client.chat.completions.create(
    model="your-model-name",
    messages=[{"role": "user", "content": "Write a short poem about APIs."}],
    stream=True,
)

for chunk in stream:
    if chunk.choices and chunk.choices[0].delta.content:
        print(chunk.choices[0].delta.content, end="", flush=True)

4.3 Python — requests (no SDK)

python
import requests

response = requests.post(
    "https://your-requrv-instance.com/api/v1/chat/completions",
    headers={
        "Authorization": "Bearer requrv_sk_YOUR_API_KEY",
        "Content-Type": "application/json",
    },
    json={
        "model": "your-model-name",
        "messages": [{"role": "user", "content": "Hello!"}],
    },
    timeout=60,
)

print(response.json()["choices"][0]["message"]["content"])

4.4 JavaScript / TypeScript — openai SDK

typescript
import OpenAI from "openai";

const client = new OpenAI({
  baseURL: "https://your-requrv-instance.com/api/v1",
  apiKey: "requrv_sk_YOUR_API_KEY",
});

const completion = await client.chat.completions.create({
  model: "your-model-name",
  messages: [
    { role: "system", content: "You are a helpful assistant." },
    { role: "user", content: "Summarize the ReQurv proxy in one sentence." },
  ],
});

console.log(completion.choices[0].message.content);

4.5 Node.js — fetch (no SDK)

javascript
const res = await fetch(
  "https://your-requrv-instance.com/api/v1/chat/completions",
  {
    method: "POST",
    headers: {
      "Content-Type": "application/json",
      Authorization: "Bearer requrv_sk_YOUR_API_KEY",
    },
    body: JSON.stringify({
      model: "your-model-name",
      messages: [{ role: "user", content: "Hello!" }],
    }),
  },
);

const data = await res.json();
console.log(data.choices[0].message.content);

4.6 LangChain (Python)

python
from langchain_openai import ChatOpenAI

llm = ChatOpenAI(
    model="your-model-name",
    base_url="https://your-requrv-instance.com/api/v1",
    api_key="requrv_sk_YOUR_API_KEY",
)

response = llm.invoke("Explain the proxy in one sentence.")
print(response.content)

4.7 LangChain (JavaScript)

javascript
import { ChatOpenAI } from "@langchain/openai";

const model = new ChatOpenAI({
  model: "your-model-name",
  configuration: {
    baseURL: "https://your-requrv-instance.com/api/v1",
    apiKey: "requrv_sk_YOUR_API_KEY",
  },
});

const response = await model.invoke("Explain the proxy in one sentence.");
console.log(response.content);

4.8 LlamaIndex

python
from llama_index.llms.openai import OpenAI

llm = OpenAI(
    model="your-model-name",
    api_base="https://your-requrv-instance.com/api/v1",
    api_key="requrv_sk_YOUR_API_KEY",
)

response = llm.complete("Explain the proxy in one sentence.")
print(response.text)

Most AI tools let you configure a custom OpenAI-compatible provider. Point them at the ReQurv base URL with your API key:

Continue.dev (config.json)
{
  "models": [
    {
      "title": "ReQurv",
      "provider": "openai",
      "model": "your-model-name",
      "apiBase": "https://your-requrv-instance.com/api/v1",
      "apiKey": "requrv_sk_YOUR_API_KEY"
    }
  ]
}
Cline (MCP settings / API config)
{
  "openAIBaseUrl": "https://your-requrv-instance.com/api/v1",
  "openAIApiKey": "requrv_sk_YOUR_API_KEY",
  "openAIModelId": "your-model-name"
}
Open WebUI (environment variables)
OPENAI_API_BASE_URL=https://your-requrv-instance.com/api/v1
OPENAI_API_KEY=requrv_sk_YOUR_API_KEY

5. Endpoints reference

All endpoints require the Authorization: Bearer <API_KEY> header. The proxy is fully OpenAI-compatible, so standard OpenAI client libraries work unchanged.

5.1 List models

GET /v1/models
GET /api/v1/models

Returns the models enabled for your organization, with their type:

Response
{
  "object": "list",
  "data": [
    {
      "id": "your-model-name",
      "object": "model",
      "created": 1721234567,
      "owned_by": "ReQurv",
      "type": "TEXT_GENERATION"
    }
  ]
}

5.2 Chat completions

POST /v1/chat/completions
POST /api/v1/chat/completions

The standard OpenAI chat completions endpoint. Supports both streaming ("stream": true) and non-streaming requests.

ParameterTypeDescription
modelstringRequired. A TEXT_GENERATION model name.
messagesarrayRequired. The conversation, with system, user, assistant, and tool roles.
prompt_idstringOptional. The ID of a configured system prompt (see 6.4).
temperaturenumberSampling temperature (default 1).
top_pnumberNucleus sampling (default 1).
max_tokens / max_completion_tokensintegerMaximum tokens to generate.
stopstring | arrayStop sequences.
presence_penalty / frequency_penaltynumberRepetition controls.
seedintegerDeterministic sampling seed.
response_formatobjectStructured output: {"type":"json_object"} or {"type":"json_schema","json_schema":{...}}; forwarded to the upstream (see 6.6).
streambooleanWhen true, returns a Server-Sent Events stream.
stream_optionsobject{ "include_usage": true } appends a final usage chunk to the stream.
extra_bodyobjectPass-through for provider-specific parameters (see 6.3).

5.3 Responses (OpenAI Responses API)

POST /v1/responses
POST /api/v1/responses

The OpenAI Responses API endpoint, for clients built on client.responses instead of client.chat.completions. Supports both streaming ("stream": true) and non-streaming requests.

ParameterTypeDescription
modelstringRequired. A TEXT_GENERATION model name.
inputstring | arrayRequired. The prompt: a plain text string, or an array of Responses API input items (messages, tool calls, tool outputs, …).
instructionsstringOptional. System-level instructions. When no prompt_id is given and the model has a system prompt, the model system prompt is used instead.
prompt_idstringOptional. The ID of a configured system prompt (see 6.4). Takes precedence over instructions.
streambooleanWhen true, returns a Server-Sent Events stream of Responses API events (response.created, response.output_text.delta, …, response.completed).
temperature / top_pnumberSampling parameters, as in chat completions.
max_output_tokensintegerMaximum tokens to generate.
tools / tool_choicearray / objectTool definitions; passed through to the upstream as-is.
previous_response_idstringOptional. Reference to a previous response for stateful conversations; passed through to the upstream as-is.
extra_bodyobjectPass-through for provider-specific parameters (see 6.3).
curl
curl https://your-requrv-instance.com/api/v1/responses \
  -H "Content-Type: application/json" \
  -H "Authorization: Bearer requrv_sk_YOUR_API_KEY" \
  -d '{
    "model": "your-model-name",
    "input": "Explain the proxy in one sentence."
  }'

The official Python SDK works unchanged:

python
from openai import OpenAI

client = OpenAI(
    base_url="https://your-requrv-instance.com/api/v1",
    api_key="requrv_sk_YOUR_API_KEY",
)

response = client.responses.create(
    model="your-model-name",
    input="Explain the proxy in one sentence.",
)
print(response.output_text)
When the model has a GUARD model, the user text in input is screened before the request is forwarded. A "controversial" verdict appends a warning to instructions (see 6.5).

5.4 Legacy text completions

POST /v1/completions
POST /api/v1/completions

The legacy completions endpoint, for clients that still use prompt instead of messages. Supports streaming.

ParameterTypeDescription
modelstringRequired. A TEXT_GENERATION model name.
promptstring | arrayRequired. The prompt text.
prompt_idstringOptional. The ID of a configured system prompt (see 6.4). Its content is prepended to the prompt text.

5.5 Embeddings

POST /v1/embeddings
POST /api/v1/embeddings

Generates vector embeddings for text. Requires an EMBEDDING model.

ParameterTypeDescription
modelstringRequired. An EMBEDDING model name.
inputstring | arrayRequired. Text to embed.
encoding_format"float" | "base64"Output format (default "float").
dimensionsintegerOptional output dimensions.

5.6 Cohere Embed v2 (multimodal)

POST /v2/embeddings
POST /api/v2/embeddings

Cohere-compatible embeddings that accept text, images, or mixed inputs. Requires an EMBEDDING model.

ParameterTypeDescription
modelstringRequired. An EMBEDDING model name.
textsarrayList of text strings to embed.
imagesarrayList of image URLs or base64 data URIs.
inputsarrayMixed content parts: { "content": [{ "type": "text", "text": "..." }] } or { "type": "image_url", "image_url": { "url": "..." } }.
input_typestringCohere input type (e.g. "search_query", "search_document").
embedding_typesarraye.g. ["float"].
truncate"END" | "START" | "NONE"Truncation strategy.

5.7 Audio transcription

POST /v1/audio/transcriptions
POST /api/v1/audio/transcriptions

Transcribes audio files to text. Requires an AUDIO_TRANSCRIPTION model. This is a multipart/form-data request.

ParameterTypeDescription
modelstringRequired. An AUDIO_TRANSCRIPTION model name.
filefileRequired. The audio file.
languagestringOptional language hint (ISO-639-1).
promptstringOptional guiding prompt.
response_formatstringjson, verbose_json, diarized_json, text, srt, or vtt.
temperaturenumberSampling temperature.
timestamp_granularitiesarray["word"], ["segment"], or both.
streambooleanWhen true, returns an SSE stream.

5.8 Speech generation (TTS)

POST /v1/audio/speech
POST /api/v1/audio/speech

Generates spoken audio from text. Requires an AUDIO_GENERATION model. This is a application/json request and the response is the raw audio file (no JSON wrapper).

Request
{
  "model": "your-tts-model",
  "input": "Ciao, benvenuto su ReQurv.",
  "voice": "alloy",
  "response_format": "mp3",
  "speed": 1.0
}
ParameterTypeDescription
modelstringRequired. An AUDIO_GENERATION model name.
inputstringRequired. The text to speak (max 4096 characters).
voicestringThe voice to use. Omit to let the upstream server use its default.
response_formatstringmp3 (default), opus, aac, flac, wav, or pcm.
speednumberPlayback speed from 0.25 to 4.0. 1.0 is the default.
instructionsstringOptional voice style instructions (ignored by some upstream servers).
stream_formatstringsse or audio: when set, the response is streamed instead of returned in one piece.

Usage is billed by minutes of generated audio, against the same audio-minutes quota as transcription.

5.9 Voice cloning

POST /v1/audio/voices
POST /api/v1/audio/voices

Clones a new voice from an audio sample. This is a multipart/form-data request. The voice is registered on the TTS engine and becomes immediately available for speech generation.

ParameterTypeDescription
namestringRequired. The voice name. Must end with a two-letter language suffix (e.g. sarah_en, ilaria_it, max_de).
audio_samplefileRequired. An .mp3 file with a sample of the voice to clone.
ref_textstringRequired. The text spoken in the audio sample.
descriptionstringOptional description of the voice.
consentbooleanRequired. Must be true: the explicit consent to the use of the voice by the person providing it.
Response
{
  "id": "sarah_en",
  "object": "voice",
  "language": "en",
  "sampleUrl": "/api/v1/voices/sarah_en/sample",
  "description": null,
  "isCloned": true
}

Errors: 400 INVALID_VOICE_NAME (missing or invalid language suffix), 400 INVALID_AUDIO_SAMPLE (not an .mp3 file), 400 REF_TEXT_REQUIRED, 400 CONSENT_REQUIRED, 404 MODEL_NOT_FOUND, 409 VOICE_ALREADY_EXISTS, 502 (upstream TTS error).

5.10 Voice deletion

DELETE /v1/audio/voices/{voice_id}
DELETE /api/v1/audio/voices/sarah_en

Permanently deletes a cloned voice from the TTS engine, the sample storage and the voice list. Voices inserted by the admin (seeding) cannot be deleted: the request returns 403 VOICE_NOT_DELETABLE.

Response
{
  "id": "sarah_en",
  "object": "voice",
  "deleted": true
}

Errors: 403 VOICE_NOT_DELETABLE, 404 VOICE_NOT_FOUND, 502 (upstream TTS error).


6. Request parameters

6.1 Common parameters

The proxy accepts the standard OpenAI sampling parameters: temperature, top_p, n, stop, max_tokens, max_completion_tokens, presence_penalty, frequency_penalty, seed, user, stream, and stream_options. Unsupported parameters are forwarded to the upstream server, which may ignore them.

6.2 System prompt fallback

If a model is configured with a default system prompt and your request does not include a system message, the proxy injects the model's system prompt automatically.

6.3 extra_body

Some model servers accept parameters that are not part of the OpenAI spec. Pass them through the extra_body object:

Request
{
  "model": "your-model-name",
  "messages": [{ "role": "user", "content": "Solve this step by step." }],
  "extra_body": {
    "chat_template_kwargs": { "enable_thinking": false }
  }
}

The values inside extra_body are merged into the request sent to the upstream model server.

6.4 System prompts via prompt_id

Manage reusable system prompts from the Prompts page in the dashboard. Each prompt has a name, an environment label (DEV or PROD), and the system prompt body. To use one in a request, pass its ID as the prompt_id parameter:

Request
{
  "model": "your-model-name",
  "messages": [{ "role": "user", "content": "Summarize this article." }],
  "prompt_id": "3f2a9c1e-...-prompt-uuid"
}

The proxy loads the prompt and injects its content as the system message. It replaces any system message you send in messages. On the legacy /v1/completions endpoint the prompt content is prepended to the prompt text instead.

If the prompt_id does not exist or does not belong to your organization, the proxy returns 404 PROMPT_NOT_FOUND.

6.5 Content guard (GUARD models)

A TEXT_GENERATION model can be associated with a GUARD model (a Qwen3Guard-style safety classifier). When configured, the proxy screens the user input before forwarding it to the target model:

  1. The proxy sends the request messages to the GUARD model.
  2. The guard classifies the input as Safe, Unsafe, or Controversial and returns the matching safety categories (e.g. Violent, PII, Jailbreak).
  3. If the input is Unsafe, the proxy immediately returns 400 CONTENT_BLOCKED with the block reason — the target model is never called.
  4. If the input is Controversial, the proxy appends a warning to the last user message and forwards the request to the target model, so the model is aware the prompt could be dangerous. The verdict is also logged.
  5. If the input is Safe, the request proceeds unchanged.
Blocked response (Unsafe)
{
  "error": "CONTENT_BLOCKED",
  "reason": "Violent",
  "categories": ["Violent"],
  "safety": "Unsafe"
}

For a Controversial verdict the last user message is rewritten as follows before being forwarded:

Rewritten user message (Controversial)
{
  "role": "user",
  "content": "Vorrei morire... \n\n\n WARNING: Content controversial, category: Suicide & Self-Harm"
}

The guard is fail-open: if the GUARD model is unavailable or errors, the request is forwarded anyway and the incident is logged, so a guard outage never takes the proxy down. Token usage of the guard model is recorded in usage/activity under the guard model name.

6.6 Structured outputs (response_format)

You can ask the model to return a structured response — most commonly, valid JSON that conforms to a schema — by sending the OpenAI-standard response_format parameter. ReQurv is a pass-through proxy: it forwards response_format (and everything in extra_body) verbatim to the upstream model server, so the constrained decoding happens on the model server, not in ReQurv. It works whenever the upstream server supports guided / structured decoding (for example vLLM with its xgrammar or guidance backend).

The two standard response_format forms are:

Request — JSON mode
{
  "model": "your-model-name",
  "messages": [{ "role": "user", "content": "Return my data as JSON." }],
  "response_format": { "type": "json_object" }
}
Request — JSON schema
{
  "model": "your-model-name",
  "messages": [{ "role": "user", "content": "Generate a JSON with the brand, model and car_type of the most iconic car from the 90s." }],
  "response_format": {
    "type": "json_schema",
    "json_schema": {
      "name": "car-description",
      "schema": {
        "type": "object",
        "properties": {
          "brand": { "type": "string" },
          "model": { "type": "string" },
          "car_type": { "type": "string", "enum": ["sedan", "SUV", "Truck", "Coupe"] }
        },
        "required": ["brand", "model", "car_type"],
        "additionalProperties": false
      }
    }
  }
}

With the openai Python SDK you can build the schema from a Pydantic model instead of writing it by hand:

python
from enum import Enum
from openai import OpenAI
from pydantic import BaseModel


class CarType(str, Enum):
    sedan = "sedan"
    suv = "SUV"
    truck = "Truck"
    coupe = "Coupe"


class CarDescription(BaseModel):
    brand: str
    model: str
    car_type: CarType


client = OpenAI(
    base_url="https://your-requrv-instance.com/api/v1",
    api_key="requrv_sk_YOUR_API_KEY",
)

response = client.chat.completions.create(
    model="your-model-name",
    messages=[{"role": "user", "content": "Generate a JSON with the brand, model and car_type of the most iconic car from the 90s."}],
    response_format={
        "type": "json_schema",
        "json_schema": {"name": "car-description", "schema": CarDescription.model_json_schema()},
    },
)

print(response.choices[0].message.content)

More constrained options (vLLM). Beyond response_format, vLLM exposes richer structured-output modes through a structured_outputs object. ReQurv merges extra_body into the request sent to the upstream model server (see 6.3), so you pass it there:

  • structured_outputs.choice — the output must be exactly one of a list of choices.
  • structured_outputs.regex — the output must match a regular expression.
  • structured_outputs.json — the output must follow a JSON schema.
  • structured_outputs.grammar — the output must follow a context-free (EBNF) grammar.
  • structured_outputs.structural_tag — a JSON schema applied to specific tags within otherwise-free text.
Request — vLLM structured_outputs via extra_body
{
  "model": "your-model-name",
  "messages": [{ "role": "user", "content": "Classify this sentiment: vLLM is wonderful!" }],
  "extra_body": {
    "structured_outputs": { "choice": ["positive", "negative"] }
  }
}
ReQurv does not validate or enforce the schema itself — response_format and extra_body reach the upstream server unchanged, so the actual enforcement (and the supported regex syntax) depends on the model server and its structured-outputs backend. The deprecated vLLM guided_json / guided_regex / guided_choice / guided_grammar fields were removed in vLLM v0.12.0 in favor of structured_outputs; prefer the structured_outputs form. The Anthropic-compatible POST /api/v1/messages endpoint forwards response_format the same way.

7. Streaming

Set "stream": true to receive the response as a Server-Sent Events (SSE) stream. Each chunk is a data: line containing a JSON object, terminated by a final data: [DONE]:

SSE stream
data: {"id":"chatcmpl-...","object":"chat.completion.chunk","choices":[{"index":0,"delta":{"role":"assistant"},"finish_reason":null}]}

data: {"id":"chatcmpl-...","object":"chat.completion.chunk","choices":[{"index":0,"delta":{"content":"Hello"},"finish_reason":null}]}

data: {"id":"chatcmpl-...","object":"chat.completion.chunk","choices":[{"index":0,"delta":{},"finish_reason":"stop"}]}

data: [DONE]

To receive a final chunk with token usage, add "stream_options": { "include_usage": true }. For Responses streams, token usage is included in the final response.completed event.

Streaming works for chat completions, legacy completions, Responses, and audio transcription. Speech generation streams with the stream_format parameter (sse or audio) instead of stream.


8. Errors and status codes

Errors are returned as JSON with an error field:

Error response
{
  "error": "MODEL_NOT_FOUND: \"unknown-model\""
}
StatusCodeMeaning
401INVALID_API_KEYThe API key is not recognized. Check the requrv_sk_ prefix and that you copied the full key.
401REVOKED_API_KEYThe API key has been revoked. Create a new one.
401EXPIRED_API_KEYThe API key has expired.
401INVALID_SESSIONNo valid session or key was provided.
400MODEL_DISABLEDThe model exists but is disabled.
400MODEL_TYPE_MISMATCHThe model exists but is not the right type for this endpoint (e.g. an EMBEDDING model sent to /chat/completions).
400CONTENT_BLOCKEDThe user input was flagged as unsafe by the model's GUARD. The response includes reason, categories, and safety.
404MODEL_NOT_FOUNDNo model with that name. List available models with GET /v1/models.
404PROMPT_NOT_FOUNDThe prompt_id does not exist or does not belong to your organization.
429CONCURRENCY_LIMIT_EXCEEDEDYour organization has too many in-flight requests. Retry after a short delay.
502UPSTREAM_ERRORThe underlying model server failed. Check the model server or try again later.

9. Model types

Models are grouped by type, and each endpoint accepts only the matching type:

TypeUsed by
TEXT_GENERATION/v1/chat/completions, /v1/completions, /v1/responses
EMBEDDING/v1/embeddings, /v2/embeddings
AUDIO_TRANSCRIPTION/v1/audio/transcriptions
AUDIO_GENERATION/v1/audio/speech
GUARDInternal safety classifier (see 6.5)

If you get a MODEL_TYPE_MISMATCH error, you are using a model of the wrong type for the endpoint.


10. Troubleshooting

"I get MODEL_NOT_FOUND." The model name must match exactly one of the models enabled for your organization. Run GET /v1/models and copy the id from the response.

"I get 401 INVALID_API_KEY." Make sure the key starts with requrv_sk_, is copied in full, and is still active. Keys are shown only once at creation — if you lost it, create a new one.

"My streaming request hangs." Streaming responses use Server-Sent Events. Make sure your client is configured to handle SSE (the official OpenAI SDKs do this automatically) and that no proxy in between buffers the response.

"The response is slow / I get 429." Your organization has a concurrency limit on simultaneous requests. Wait for in-flight requests to finish, or reduce the number of parallel calls.

"I get 502 UPSTREAM_ERROR." The model server behind the proxy is unreachable or returned an error. This usually means the underlying model is down or misconfigured — contact your administrator.


11. Next steps

  • Inspect the full OpenAPI specification at https://your-requrv-instance.com/api/openapi/json.
  • Track token usage and activity in the ReQurv Dashboard.
  • Manage and rotate keys in the API Keys page.