> For clean Markdown of any page, append .md to the page URL.
> For a complete documentation index, see https://api.qdrant.tech/llms.txt.
> For AI client integration (Claude Code, Cursor, etc.), connect to the MCP server at https://api.qdrant.tech/_mcp/server.

# Get global quotas

GET http://localhost:6333/quotas

Get the cluster-wide resource quota configuration, together with the current utilization it is measured against.
The configuration is the same on every peer, but the reported utilization is for the node serving this request only -
memory and disk are node-local, so query each peer to see where the whole cluster stands.


Reference: https://api.qdrant.tech/api-reference/quotas/get-quotas

## Authentication

- `api-key` header (required) — API Key authentication via header

## Servers

- `http://localhost:6333` (http, default)
- `https://localhost:6333` (https)

## Response

### 200

successful operation

- `usage` (QuotasGetResponsesContentApplicationJsonSchemaUsage, optional)
- `time` (double, optional) — Time spent to process this request
- `status` (string, optional)
- `result` (QuotaStatus, optional) — Quota configuration in effect, and how close each peer is to it. The configuration is cluster-wide; the utilization is not. `usage` is the node that served the request, and `peers` is what every peer that answered reports about itself — memory and disk are node-local, so one peer being under its limit says nothing about the others.

## Types

### QuotasGetResponsesContentApplicationJsonSchemaUsage

### QuotaStatus

Quota configuration in effect, and how close each peer is to it. The configuration is cluster-wide; the utilization is not. `usage` is the node that served the request, and `peers` is what every peer that answered reports about itself — memory and disk are node-local, so one peer being under its limit says nothing about the others.

- `config` (QuotaConfig, required) — Cluster-wide limits on node resources. An unset limit means the corresponding resource is not capped. Limits are only enforced while `enabled` is true.
- `usage` (QuotaUsage, required) — Utilization of the quota-managed resources **on this node alone** — memory and disk are node-local, so a peer under its limit says nothing about the rest of the cluster. A field is `null` when the platform does not expose the underlying stat.
- `peers` (map from string to PeerQuotaUsage, optional, nullable) — Utilization reported by each peer, keyed by peer ID, including the one that served the request. Only peers that answered are listed: a peer missing from the map could not be reached, which is itself worth seeing. Absent entirely outside distributed mode, where there are no peers to ask.

### Usage

Usage of the hardware resources, spent to process the request

- `hardware` (UsageHardware, optional)
- `inference` (UsageInference, optional)

### QuotaConfig

Cluster-wide limits on node resources. An unset limit means the corresponding resource is not capped. Limits are only enforced while `enabled` is true.

- `enabled` (boolean, optional, default: false) — Whether the limits below are enforced.
- `max_resident_memory_percent` (integer, optional, nullable) — Reject memory-consuming updates once process resident memory reaches this percentage of total system memory (or of the cgroup limit, if one applies).
- `max_disk_usage_percent` (integer, optional, nullable) — Reject disk-consuming updates once the filesystem hosting the storage directory is filled to this percentage of its capacity.
- `release_margin_percent` (integer, optional, nullable) — How many percentage points below its limit a resource has to fall before this node starts accepting work again. Without a margin, a resource resting on its limit crosses it in both directions on the noise between two readings, putting the node in and out of service each time — and restarting a shard recovery with it. Raise it where usage is volatile; `0` disables the margin and releases as soon as usage is back under the limit. Unset leaves the built-in default in force, so a config written today does not pin a number that a later release may want to revise.

### QuotaUsage

Utilization of the quota-managed resources **on this node alone** — memory and disk are node-local, so a peer under its limit says nothing about the rest of the cluster. A field is `null` when the platform does not expose the underlying stat.

- `resident_memory_percent` (integer, optional, nullable) — Resident memory of this node's process, as a percentage of the memory available to it (cgroup limit if one applies, else total system memory).
- `disk_usage_percent` (integer, optional, nullable) — Used space of this node's storage filesystem, as a percentage of its capacity.

### PeerQuotaUsage

What one peer reports about the quota it is enforcing.

- `exceeded` (boolean, required) — Whether this peer is at or over one of the enforced limits, and so is currently refusing updates. Always false while the quota is disabled.
- `resident_memory_percent` (integer, optional, nullable) — Resident memory of this node's process, as a percentage of the memory available to it (cgroup limit if one applies, else total system memory).
- `disk_usage_percent` (integer, optional, nullable) — Used space of this node's storage filesystem, as a percentage of its capacity.

### UsageHardware

### UsageInference

### HardwareUsage

Usage of the hardware resources, spent to process the request

- `cpu` (integer, required)
- `payload_io_read` (integer, required)
- `payload_io_write` (integer, required)
- `payload_index_io_read` (integer, required)
- `payload_index_io_write` (integer, required)
- `vector_io_read` (integer, required)
- `vector_io_write` (integer, required)

### InferenceUsage

- `models` (map from string to ModelUsage, required)

### ModelUsage

- `tokens` (uint64, required)

## Examples

**Response**

```json
{
  "usage": {
    "hardware": {
      "cpu": 1,
      "payload_io_read": 1,
      "payload_io_write": 1,
      "payload_index_io_read": 1,
      "payload_index_io_write": 1,
      "vector_io_read": 1,
      "vector_io_write": 1
    },
    "inference": {
      "models": {}
    }
  },
  "time": 0.002,
  "status": "ok",
  "result": {
    "config": {
      "enabled": false,
      "max_resident_memory_percent": 1,
      "max_disk_usage_percent": 1,
      "release_margin_percent": 1
    },
    "usage": {
      "resident_memory_percent": 1,
      "disk_usage_percent": 1
    },
    "peers": {}
  }
}
```

**SDK Code**

```python
import requests

url = "http://localhost:6333/quotas"

headers = {"api-key": "<apiKey>"}

response = requests.get(url, headers=headers)

print(response.json())
```

```javascript
const url = 'http://localhost:6333/quotas';
const options = {method: 'GET', headers: {'api-key': '<apiKey>'}};

try {
  const response = await fetch(url, options);
  const data = await response.json();
  console.log(data);
} catch (error) {
  console.error(error);
}
```

```go
package main

import (
	"fmt"
	"net/http"
	"io"
)

func main() {

	url := "http://localhost:6333/quotas"

	req, _ := http.NewRequest("GET", url, nil)

	req.Header.Add("api-key", "<apiKey>")

	res, _ := http.DefaultClient.Do(req)

	defer res.Body.Close()
	body, _ := io.ReadAll(res.Body)

	fmt.Println(res)
	fmt.Println(string(body))

}
```

```ruby
require 'uri'
require 'net/http'

url = URI("http://localhost:6333/quotas")

http = Net::HTTP.new(url.host, url.port)

request = Net::HTTP::Get.new(url)
request["api-key"] = '<apiKey>'

response = http.request(request)
puts response.read_body
```

```java
import com.mashape.unirest.http.HttpResponse;
import com.mashape.unirest.http.Unirest;

HttpResponse<String> response = Unirest.get("http://localhost:6333/quotas")
  .header("api-key", "<apiKey>")
  .asString();
```

```php
<?php
require_once('vendor/autoload.php');

$client = new \GuzzleHttp\Client();

$response = $client->request('GET', 'http://localhost:6333/quotas', [
  'headers' => [
    'api-key' => '<apiKey>',
  ],
]);

echo $response->getBody();
```

```csharp
using RestSharp;

var client = new RestClient("http://localhost:6333/quotas");
var request = new RestRequest(Method.GET);
request.AddHeader("api-key", "<apiKey>");
IRestResponse response = client.Execute(request);
```

```swift
import Foundation

let headers = ["api-key": "<apiKey>"]

let request = NSMutableURLRequest(url: NSURL(string: "http://localhost:6333/quotas")! as URL,
                                        cachePolicy: .useProtocolCachePolicy,
                                    timeoutInterval: 10.0)
request.httpMethod = "GET"
request.allHTTPHeaderFields = headers

let session = URLSession.shared
let dataTask = session.dataTask(with: request as URLRequest, completionHandler: { (data, response, error) -> Void in
  if (error != nil) {
    print(error as Any)
  } else {
    let httpResponse = response as? HTTPURLResponse
    print(httpResponse)
  }
})

dataTask.resume()
```