Documentazione
Tutto quello che serve per connettere la tua applicazione all'API ReQurv Hive, compatibile con OpenAI.
Dettagli di connessione
Usa questi valori in qualsiasi client compatibile con OpenAI.
URL base
https://hive.requrv.ai/api/v1
Intestazione di autenticazione
Authorization: Bearer requrv_sk_…
ReQurv Proxy — OpenAI-Compatible API
ReQurv exposes your organization's LLM models through a drop-in, OpenAI-compatible API. If your tool, framework, or custom code can talk to the OpenAI API, it can talk to ReQurv — you only need to change the base URL and the API key.
The proxy supports chat completions, legacy text completions, embeddings, audio transcription, and speech generation, with both synchronous and streaming responses.
1. Prerequisites
Before you can make your first request you need an API key.
- Open the API Keys page from the sidebar.
- Click Crea chiave (Create key) and give the key a name.
- Copy the secret immediately — it is shown only once.
API keys always start with requrv_sk_ and are sent in the Authorization header:
Authorization: Bearer requrv_sk_xxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxx2. Connection details
| Setting | Value |
|---|---|
| Base URL | https://your-requrv-instance.com/api/v1 |
| Authentication | Authorization: Bearer <YOUR_API_KEY> |
| Models | Listed by GET /v1/models (see below) |
Replace https://your-requrv-instance.com with the address of your ReQurv instance (the same URL you use to open the dashboard). The OpenAI-compatible endpoints live under the /api/v1 path.
The model name you send in requests must match one of the models configured for your organization. To see the exact names available to you:
curl https://your-requrv-instance.com/api/v1/models \
-H "Authorization: Bearer requrv_sk_YOUR_API_KEY"3. Quick start
The simplest possible chat completion with curl:
curl https://your-requrv-instance.com/api/v1/chat/completions \
-H "Content-Type: application/json" \
-H "Authorization: Bearer requrv_sk_YOUR_API_KEY" \
-d '{
"model": "your-model-name",
"messages": [
{ "role": "system", "content": "You are a helpful assistant." },
{ "role": "user", "content": "Explain the proxy in one sentence." }
]
}'A successful response looks like this:
{
"id": "chatcmpl-...",
"object": "chat.completion",
"created": 1721234567,
"model": "your-model-name",
"choices": [
{
"index": 0,
"message": {
"role": "assistant",
"content": "The ReQurv proxy is an OpenAI-compatible gateway that lets you call your organization's LLM models from any tool or framework."
},
"finish_reason": "stop"
}
],
"usage": {
"prompt_tokens": 24,
"completion_tokens": 22,
"total_tokens": 46
}
}4. Code examples
All examples use https://your-requrv-instance.com/api/v1 as the base URL and requrv_sk_YOUR_API_KEY as the key. Replace both with your own values.
4.1 Python — openai SDK
from openai import OpenAI
client = OpenAI(
base_url="https://your-requrv-instance.com/api/v1",
api_key="requrv_sk_YOUR_API_KEY",
)
response = client.chat.completions.create(
model="your-model-name",
messages=[
{"role": "system", "content": "You are a concise assistant."},
{"role": "user", "content": "What is the capital of France?"},
],
)
print(response.choices[0].message.content)4.2 Python — streaming with the openai SDK
from openai import OpenAI
client = OpenAI(
base_url="https://your-requrv-instance.com/api/v1",
api_key="requrv_sk_YOUR_API_KEY",
)
stream = client.chat.completions.create(
model="your-model-name",
messages=[{"role": "user", "content": "Write a short poem about APIs."}],
stream=True,
)
for chunk in stream:
if chunk.choices and chunk.choices[0].delta.content:
print(chunk.choices[0].delta.content, end="", flush=True)4.3 Python — requests (no SDK)
import requests
response = requests.post(
"https://your-requrv-instance.com/api/v1/chat/completions",
headers={
"Authorization": "Bearer requrv_sk_YOUR_API_KEY",
"Content-Type": "application/json",
},
json={
"model": "your-model-name",
"messages": [{"role": "user", "content": "Hello!"}],
},
timeout=60,
)
print(response.json()["choices"][0]["message"]["content"])4.4 JavaScript / TypeScript — openai SDK
import OpenAI from "openai";
const client = new OpenAI({
baseURL: "https://your-requrv-instance.com/api/v1",
apiKey: "requrv_sk_YOUR_API_KEY",
});
const completion = await client.chat.completions.create({
model: "your-model-name",
messages: [
{ role: "system", content: "You are a helpful assistant." },
{ role: "user", content: "Summarize the ReQurv proxy in one sentence." },
],
});
console.log(completion.choices[0].message.content);4.5 Node.js — fetch (no SDK)
const res = await fetch(
"https://your-requrv-instance.com/api/v1/chat/completions",
{
method: "POST",
headers: {
"Content-Type": "application/json",
Authorization: "Bearer requrv_sk_YOUR_API_KEY",
},
body: JSON.stringify({
model: "your-model-name",
messages: [{ role: "user", content: "Hello!" }],
}),
},
);
const data = await res.json();
console.log(data.choices[0].message.content);4.6 LangChain (Python)
from langchain_openai import ChatOpenAI
llm = ChatOpenAI(
model="your-model-name",
base_url="https://your-requrv-instance.com/api/v1",
api_key="requrv_sk_YOUR_API_KEY",
)
response = llm.invoke("Explain the proxy in one sentence.")
print(response.content)4.7 LangChain (JavaScript)
import { ChatOpenAI } from "@langchain/openai";
const model = new ChatOpenAI({
model: "your-model-name",
configuration: {
baseURL: "https://your-requrv-instance.com/api/v1",
apiKey: "requrv_sk_YOUR_API_KEY",
},
});
const response = await model.invoke("Explain the proxy in one sentence.");
console.log(response.content);4.8 LlamaIndex
from llama_index.llms.openai import OpenAI
llm = OpenAI(
model="your-model-name",
api_base="https://your-requrv-instance.com/api/v1",
api_key="requrv_sk_YOUR_API_KEY",
)
response = llm.complete("Explain the proxy in one sentence.")
print(response.text)4.9 Popular tools
Most AI tools let you configure a custom OpenAI-compatible provider. Point them at the ReQurv base URL with your API key:
{
"models": [
{
"title": "ReQurv",
"provider": "openai",
"model": "your-model-name",
"apiBase": "https://your-requrv-instance.com/api/v1",
"apiKey": "requrv_sk_YOUR_API_KEY"
}
]
}{
"openAIBaseUrl": "https://your-requrv-instance.com/api/v1",
"openAIApiKey": "requrv_sk_YOUR_API_KEY",
"openAIModelId": "your-model-name"
}OPENAI_API_BASE_URL=https://your-requrv-instance.com/api/v1
OPENAI_API_KEY=requrv_sk_YOUR_API_KEY5. Endpoints reference
All endpoints require the Authorization: Bearer <API_KEY> header. The proxy is fully OpenAI-compatible, so standard OpenAI client libraries work unchanged.
5.1 List models
GET /api/v1/modelsReturns the models enabled for your organization, with their type:
{
"object": "list",
"data": [
{
"id": "your-model-name",
"object": "model",
"created": 1721234567,
"owned_by": "ReQurv",
"type": "TEXT_GENERATION"
}
]
}5.2 Chat completions
POST /api/v1/chat/completionsThe standard OpenAI chat completions endpoint. Supports both streaming ("stream": true) and non-streaming requests.
| Parameter | Type | Description |
|---|---|---|
model | string | Required. A TEXT_GENERATION model name. |
messages | array | Required. The conversation, with system, user, assistant, and tool roles. |
prompt_id | string | Optional. The ID of a configured system prompt (see 6.4). |
temperature | number | Sampling temperature (default 1). |
top_p | number | Nucleus sampling (default 1). |
max_tokens / max_completion_tokens | integer | Maximum tokens to generate. |
stop | string | array | Stop sequences. |
presence_penalty / frequency_penalty | number | Repetition controls. |
seed | integer | Deterministic sampling seed. |
response_format | object | Structured output: {"type":"json_object"} or {"type":"json_schema","json_schema":{...}}; forwarded to the upstream (see 6.6). |
stream | boolean | When true, returns a Server-Sent Events stream. |
stream_options | object | { "include_usage": true } appends a final usage chunk to the stream. |
extra_body | object | Pass-through for provider-specific parameters (see 6.3). |
5.3 Responses (OpenAI Responses API)
POST /api/v1/responsesThe OpenAI Responses API endpoint, for clients built on client.responses instead of client.chat.completions. Supports both streaming ("stream": true) and non-streaming requests.
| Parameter | Type | Description |
|---|---|---|
model | string | Required. A TEXT_GENERATION model name. |
input | string | array | Required. The prompt: a plain text string, or an array of Responses API input items (messages, tool calls, tool outputs, …). |
instructions | string | Optional. System-level instructions. When no prompt_id is given and the model has a system prompt, the model system prompt is used instead. |
prompt_id | string | Optional. The ID of a configured system prompt (see 6.4). Takes precedence over instructions. |
stream | boolean | When true, returns a Server-Sent Events stream of Responses API events (response.created, response.output_text.delta, …, response.completed). |
temperature / top_p | number | Sampling parameters, as in chat completions. |
max_output_tokens | integer | Maximum tokens to generate. |
tools / tool_choice | array / object | Tool definitions; passed through to the upstream as-is. |
previous_response_id | string | Optional. Reference to a previous response for stateful conversations; passed through to the upstream as-is. |
extra_body | object | Pass-through for provider-specific parameters (see 6.3). |
curl https://your-requrv-instance.com/api/v1/responses \
-H "Content-Type: application/json" \
-H "Authorization: Bearer requrv_sk_YOUR_API_KEY" \
-d '{
"model": "your-model-name",
"input": "Explain the proxy in one sentence."
}'The official Python SDK works unchanged:
from openai import OpenAI
client = OpenAI(
base_url="https://your-requrv-instance.com/api/v1",
api_key="requrv_sk_YOUR_API_KEY",
)
response = client.responses.create(
model="your-model-name",
input="Explain the proxy in one sentence.",
)
print(response.output_text)GUARD model, the user text in input is screened before the request is forwarded. A "controversial" verdict appends a warning to instructions (see 6.5).5.4 Legacy text completions
POST /api/v1/completionsThe legacy completions endpoint, for clients that still use prompt instead of messages. Supports streaming.
| Parameter | Type | Description |
|---|---|---|
model | string | Required. A TEXT_GENERATION model name. |
prompt | string | array | Required. The prompt text. |
prompt_id | string | Optional. The ID of a configured system prompt (see 6.4). Its content is prepended to the prompt text. |
5.5 Embeddings
POST /api/v1/embeddingsGenerates vector embeddings for text. Requires an EMBEDDING model.
| Parameter | Type | Description |
|---|---|---|
model | string | Required. An EMBEDDING model name. |
input | string | array | Required. Text to embed. |
encoding_format | "float" | "base64" | Output format (default "float"). |
dimensions | integer | Optional output dimensions. |
5.6 Cohere Embed v2 (multimodal)
POST /api/v2/embeddingsCohere-compatible embeddings that accept text, images, or mixed inputs. Requires an EMBEDDING model.
| Parameter | Type | Description |
|---|---|---|
model | string | Required. An EMBEDDING model name. |
texts | array | List of text strings to embed. |
images | array | List of image URLs or base64 data URIs. |
inputs | array | Mixed content parts: { "content": [{ "type": "text", "text": "..." }] } or { "type": "image_url", "image_url": { "url": "..." } }. |
input_type | string | Cohere input type (e.g. "search_query", "search_document"). |
embedding_types | array | e.g. ["float"]. |
truncate | "END" | "START" | "NONE" | Truncation strategy. |
5.7 Audio transcription
POST /api/v1/audio/transcriptionsTranscribes audio files to text. Requires an AUDIO_TRANSCRIPTION model. This is a multipart/form-data request.
| Parameter | Type | Description |
|---|---|---|
model | string | Required. An AUDIO_TRANSCRIPTION model name. |
file | file | Required. The audio file. |
language | string | Optional language hint (ISO-639-1). |
prompt | string | Optional guiding prompt. |
response_format | string | json, verbose_json, diarized_json, text, srt, or vtt. |
temperature | number | Sampling temperature. |
timestamp_granularities | array | ["word"], ["segment"], or both. |
stream | boolean | When true, returns an SSE stream. |
5.8 Speech generation (TTS)
POST /api/v1/audio/speechGenerates spoken audio from text. Requires an AUDIO_GENERATION model. This is a application/json request and the response is the raw audio file (no JSON wrapper).
{
"model": "your-tts-model",
"input": "Ciao, benvenuto su ReQurv.",
"voice": "alloy",
"response_format": "mp3",
"speed": 1.0
}| Parameter | Type | Description |
|---|---|---|
model | string | Required. An AUDIO_GENERATION model name. |
input | string | Required. The text to speak (max 4096 characters). |
voice | string | The voice to use. Omit to let the upstream server use its default. |
response_format | string | mp3 (default), opus, aac, flac, wav, or pcm. |
speed | number | Playback speed from 0.25 to 4.0. 1.0 is the default. |
instructions | string | Optional voice style instructions (ignored by some upstream servers). |
stream_format | string | sse or audio: when set, the response is streamed instead of returned in one piece. |
Usage is billed by minutes of generated audio, against the same audio-minutes quota as transcription.
5.9 Voice cloning
POST /api/v1/audio/voicesClones a new voice from an audio sample. This is a multipart/form-data request. The voice is registered on the TTS engine and becomes immediately available for speech generation.
| Parameter | Type | Description |
|---|---|---|
name | string | Required. The voice name. Must end with a two-letter language suffix (e.g. sarah_en, ilaria_it, max_de). |
audio_sample | file | Required. An .mp3 file with a sample of the voice to clone. |
ref_text | string | Required. The text spoken in the audio sample. |
description | string | Optional description of the voice. |
consent | boolean | Required. Must be true: the explicit consent to the use of the voice by the person providing it. |
{
"id": "sarah_en",
"object": "voice",
"language": "en",
"sampleUrl": "/api/v1/voices/sarah_en/sample",
"description": null,
"isCloned": true
}Errors: 400 INVALID_VOICE_NAME (missing or invalid language suffix), 400 INVALID_AUDIO_SAMPLE (not an .mp3 file), 400 REF_TEXT_REQUIRED, 400 CONSENT_REQUIRED, 404 MODEL_NOT_FOUND, 409 VOICE_ALREADY_EXISTS, 502 (upstream TTS error).
5.10 Voice deletion
DELETE /api/v1/audio/voices/sarah_enPermanently deletes a cloned voice from the TTS engine, the sample storage and the voice list. Voices inserted by the admin (seeding) cannot be deleted: the request returns 403 VOICE_NOT_DELETABLE.
{
"id": "sarah_en",
"object": "voice",
"deleted": true
}Errors: 403 VOICE_NOT_DELETABLE, 404 VOICE_NOT_FOUND, 502 (upstream TTS error).
6. Request parameters
6.1 Common parameters
The proxy accepts the standard OpenAI sampling parameters: temperature, top_p, n, stop, max_tokens, max_completion_tokens, presence_penalty, frequency_penalty, seed, user, stream, and stream_options. Unsupported parameters are forwarded to the upstream server, which may ignore them.
6.2 System prompt fallback
If a model is configured with a default system prompt and your request does not include a system message, the proxy injects the model's system prompt automatically.
6.3 extra_body
Some model servers accept parameters that are not part of the OpenAI spec. Pass them through the extra_body object:
{
"model": "your-model-name",
"messages": [{ "role": "user", "content": "Solve this step by step." }],
"extra_body": {
"chat_template_kwargs": { "enable_thinking": false }
}
}The values inside extra_body are merged into the request sent to the upstream model server.
6.4 System prompts via prompt_id
Manage reusable system prompts from the Prompts page in the dashboard. Each prompt has a name, an environment label (DEV or PROD), and the system prompt body. To use one in a request, pass its ID as the prompt_id parameter:
{
"model": "your-model-name",
"messages": [{ "role": "user", "content": "Summarize this article." }],
"prompt_id": "3f2a9c1e-...-prompt-uuid"
}The proxy loads the prompt and injects its content as the system message. It replaces any system message you send in messages. On the legacy /v1/completions endpoint the prompt content is prepended to the prompt text instead.
If the prompt_id does not exist or does not belong to your organization, the proxy returns 404 PROMPT_NOT_FOUND.
6.5 Content guard (GUARD models)
A TEXT_GENERATION model can be associated with a GUARD model (a Qwen3Guard-style safety classifier). When configured, the proxy screens the user input before forwarding it to the target model:
- The proxy sends the request messages to the
GUARDmodel. - The guard classifies the input as
Safe,Unsafe, orControversialand returns the matching safety categories (e.g.Violent,PII,Jailbreak). - If the input is
Unsafe, the proxy immediately returns400 CONTENT_BLOCKEDwith the block reason — the target model is never called. - If the input is
Controversial, the proxy appends a warning to the last user message and forwards the request to the target model, so the model is aware the prompt could be dangerous. The verdict is also logged. - If the input is
Safe, the request proceeds unchanged.
{
"error": "CONTENT_BLOCKED",
"reason": "Violent",
"categories": ["Violent"],
"safety": "Unsafe"
}For a Controversial verdict the last user message is rewritten as follows before being forwarded:
{
"role": "user",
"content": "Vorrei morire... \n\n\n WARNING: Content controversial, category: Suicide & Self-Harm"
}The guard is fail-open: if the GUARD model is unavailable or errors, the request is forwarded anyway and the incident is logged, so a guard outage never takes the proxy down. Token usage of the guard model is recorded in usage/activity under the guard model name.
6.6 Structured outputs (response_format)
You can ask the model to return a structured response — most commonly, valid JSON that conforms to a schema — by sending the OpenAI-standard response_format parameter. ReQurv is a pass-through proxy: it forwards response_format (and everything in extra_body) verbatim to the upstream model server, so the constrained decoding happens on the model server, not in ReQurv. It works whenever the upstream server supports guided / structured decoding (for example vLLM with its xgrammar or guidance backend).
The two standard response_format forms are:
{
"model": "your-model-name",
"messages": [{ "role": "user", "content": "Return my data as JSON." }],
"response_format": { "type": "json_object" }
}{
"model": "your-model-name",
"messages": [{ "role": "user", "content": "Generate a JSON with the brand, model and car_type of the most iconic car from the 90s." }],
"response_format": {
"type": "json_schema",
"json_schema": {
"name": "car-description",
"schema": {
"type": "object",
"properties": {
"brand": { "type": "string" },
"model": { "type": "string" },
"car_type": { "type": "string", "enum": ["sedan", "SUV", "Truck", "Coupe"] }
},
"required": ["brand", "model", "car_type"],
"additionalProperties": false
}
}
}
}With the openai Python SDK you can build the schema from a Pydantic model instead of writing it by hand:
from enum import Enum
from openai import OpenAI
from pydantic import BaseModel
class CarType(str, Enum):
sedan = "sedan"
suv = "SUV"
truck = "Truck"
coupe = "Coupe"
class CarDescription(BaseModel):
brand: str
model: str
car_type: CarType
client = OpenAI(
base_url="https://your-requrv-instance.com/api/v1",
api_key="requrv_sk_YOUR_API_KEY",
)
response = client.chat.completions.create(
model="your-model-name",
messages=[{"role": "user", "content": "Generate a JSON with the brand, model and car_type of the most iconic car from the 90s."}],
response_format={
"type": "json_schema",
"json_schema": {"name": "car-description", "schema": CarDescription.model_json_schema()},
},
)
print(response.choices[0].message.content)More constrained options (vLLM). Beyond response_format, vLLM exposes richer structured-output modes through a structured_outputs object. ReQurv merges extra_body into the request sent to the upstream model server (see 6.3), so you pass it there:
structured_outputs.choice— the output must be exactly one of a list of choices.structured_outputs.regex— the output must match a regular expression.structured_outputs.json— the output must follow a JSON schema.structured_outputs.grammar— the output must follow a context-free (EBNF) grammar.structured_outputs.structural_tag— a JSON schema applied to specific tags within otherwise-free text.
{
"model": "your-model-name",
"messages": [{ "role": "user", "content": "Classify this sentiment: vLLM is wonderful!" }],
"extra_body": {
"structured_outputs": { "choice": ["positive", "negative"] }
}
}response_format and extra_body reach the upstream server unchanged, so the actual enforcement (and the supported regex syntax) depends on the model server and its structured-outputs backend. The deprecated vLLM guided_json / guided_regex / guided_choice / guided_grammar fields were removed in vLLM v0.12.0 in favor of structured_outputs; prefer the structured_outputs form. The Anthropic-compatible POST /api/v1/messages endpoint forwards response_format the same way.7. Streaming
Set "stream": true to receive the response as a Server-Sent Events (SSE) stream. Each chunk is a data: line containing a JSON object, terminated by a final data: [DONE]:
data: {"id":"chatcmpl-...","object":"chat.completion.chunk","choices":[{"index":0,"delta":{"role":"assistant"},"finish_reason":null}]}
data: {"id":"chatcmpl-...","object":"chat.completion.chunk","choices":[{"index":0,"delta":{"content":"Hello"},"finish_reason":null}]}
data: {"id":"chatcmpl-...","object":"chat.completion.chunk","choices":[{"index":0,"delta":{},"finish_reason":"stop"}]}
data: [DONE]To receive a final chunk with token usage, add "stream_options": { "include_usage": true }. For Responses streams, token usage is included in the final response.completed event.
Streaming works for chat completions, legacy completions, Responses, and audio transcription. Speech generation streams with the stream_format parameter (sse or audio) instead of stream.
8. Errors and status codes
Errors are returned as JSON with an error field:
{
"error": "MODEL_NOT_FOUND: \"unknown-model\""
}| Status | Code | Meaning |
|---|---|---|
401 | INVALID_API_KEY | The API key is not recognized. Check the requrv_sk_ prefix and that you copied the full key. |
401 | REVOKED_API_KEY | The API key has been revoked. Create a new one. |
401 | EXPIRED_API_KEY | The API key has expired. |
401 | INVALID_SESSION | No valid session or key was provided. |
400 | MODEL_DISABLED | The model exists but is disabled. |
400 | MODEL_TYPE_MISMATCH | The model exists but is not the right type for this endpoint (e.g. an EMBEDDING model sent to /chat/completions). |
400 | CONTENT_BLOCKED | The user input was flagged as unsafe by the model's GUARD. The response includes reason, categories, and safety. |
404 | MODEL_NOT_FOUND | No model with that name. List available models with GET /v1/models. |
404 | PROMPT_NOT_FOUND | The prompt_id does not exist or does not belong to your organization. |
429 | CONCURRENCY_LIMIT_EXCEEDED | Your organization has too many in-flight requests. Retry after a short delay. |
502 | UPSTREAM_ERROR | The underlying model server failed. Check the model server or try again later. |
9. Model types
Models are grouped by type, and each endpoint accepts only the matching type:
| Type | Used by |
|---|---|
TEXT_GENERATION | /v1/chat/completions, /v1/completions, /v1/responses |
EMBEDDING | /v1/embeddings, /v2/embeddings |
AUDIO_TRANSCRIPTION | /v1/audio/transcriptions |
AUDIO_GENERATION | /v1/audio/speech |
GUARD | Internal safety classifier (see 6.5) |
If you get a MODEL_TYPE_MISMATCH error, you are using a model of the wrong type for the endpoint.
10. Troubleshooting
"I get MODEL_NOT_FOUND."
The model name must match exactly one of the models enabled for your organization. Run GET /v1/models and copy the id from the response.
"I get 401 INVALID_API_KEY."
Make sure the key starts with requrv_sk_, is copied in full, and is still active. Keys are shown only once at creation — if you lost it, create a new one.
"My streaming request hangs." Streaming responses use Server-Sent Events. Make sure your client is configured to handle SSE (the official OpenAI SDKs do this automatically) and that no proxy in between buffers the response.
"The response is slow / I get 429."
Your organization has a concurrency limit on simultaneous requests. Wait for in-flight requests to finish, or reduce the number of parallel calls.
"I get 502 UPSTREAM_ERROR."
The model server behind the proxy is unreachable or returned an error. This usually means the underlying model is down or misconfigured — contact your administrator.
11. Next steps
- Inspect the full OpenAPI specification at
https://your-requrv-instance.com/api/openapi/json. - Track token usage and activity in the ReQurv Dashboard.
- Manage and rotate keys in the API Keys page.