> ## Documentation Index
> Fetch the complete documentation index at: https://docs.zylon.ai/llms.txt
> Use this file to discover all available pages before exploring further.

# Hardware requirements

Zylon only has one strict requirement regarding hardware: **it must have access to a GPU with NVIDIA CUDA capabilities and enough VRAM to run the `baseline-96g` preset tier**. In practice, that means planning for **80GB to 96GB of VRAM or more**, depending on the GPU configuration. Depending on the hardware, some AI models might be restricted, so to ensure compatibility aim for newest compatible CUDA versions (12.6+).

Our recommended specifications for the best experience are:

* Operative System: Ubuntu **SERVER** 22.04 LTS or 24.04 LTS, clean installation, without additional drivers or software installed.
* CPU: Minimum of 12 core CPU amd64 architecture (x86\_64), but 16 cores are recommended
* RAM: Minimum of 64GB, but 128GB are recommended
* Storage: Unbounded, minimum of 4TB of fast storage, but 8TB are recommended
* GPU:
  * Bare metal
    * Server GPU (external cooling required ⚠️)
      * NVIDIA H200 (141 GB)
      * NVIDIA RTX PRO 6000 Server Edition (96 GB)
      * NVIDIA H100 (80 / 96 GB)
      * NVIDIA A100 (80 GB)
    * Workstation GPU (embedded cooling system)
      * NVIDIA RTX PRO 6000 (Workstation) (96 GB)
  * AWS
    * g6e.xlarge [https://instances.vantage.sh/aws/ec2/g6e.xlarge](https://instances.vantage.sh/aws/ec2/g6e.xlarge)
  * Azure Cloud
    * Choose an Azure GPU SKU that provides enough total VRAM for the `baseline-96g` preset tier (typically 80GB to 96GB total VRAM or more).
    * NVv3 (Requires quantized models) [https://learn.microsoft.com/en-us/azure/virtual-machines/sizes/gpu-accelerated/nvv3-series?tabs=sizebasic](https://learn.microsoft.com/en-us/azure/virtual-machines/sizes/gpu-accelerated/nvv3-series?tabs=sizebasic)

## What GPU should I buy?

Finding the right GPU for your system can be a tricky process. For example, two GPUs with the same vRAM might not perform the same:

* H100 (server) and A100 80GB are representative options for the `baseline-96g` preset tier.

On the other hand, depending on the GPU you chose, other features that might impact AI quality will be enabled, if any of them is relevant for your uses cases factor that in for your decision:

|                        | A100 / H100 |
| ---------------------- | ----------- |
| Requires Workstation   | ✅           |
| LLM                    | ✅           |
| Reranker\*             | ✅           |
| Multi-model (images)\* | ✅           |

\*Currently in development/under testing — subject to change in the future.

For **on-premise bare metal environments** (the usual scenario for Zylon clients), an important factor would be your ability to properly cool the GPU installed in your machine. If you don’t want to take care of it or lack the experience, go for a desktop hardware option. But keep in mind that in case you want to run bigger models or provide service to several hundreds of users, you might need to install a rack with a couple in parallel or be forced to move to server hardware models.

Another important factor would be the investment, specially regarding the GPU. The price ranges May 5, 2025 for the aforementioned models are:

| **GPU Model**                     | **vRAM (GB)** | **Price (USD)**   |
| --------------------------------- | ------------- | ----------------- |
| NVIDIA RTX PRO 6000 (Workstation) | 96            | TBD               |
| NVIDIA H100 (PCIe)                | 80            | $25,000 – $30,000 |
| NVIDIA H100 (SXM)                 | 80            | $35,000 – $40,000 |
| NVIDIA H100 (NVL)                 | 96            | $40,000 - $45,000 |
| NVIDIA A100 (PCIe)                | 80            | $18,000 - $20,000 |
| NVIDIA A100 (SXM)                 | 80            | $20,000 - $25,000 |
| NVIDIA H200                       | 141           | $30,000 - $32,000 |

In any case, as a direct answer to the question of which GPU should you buy, our current baseline recommendation is a configuration that supports the **`baseline-96g` preset tier**, which in practice typically means **80GB to 96GB of total VRAM or more**.

## Reference hardware for mid-size organization

If you need to acquired your AI-capable equipment from scratch, as of July 29, 2025 please consider the following [hardware recommendation](https://pcpartpicker.com/user/pablo.zylon/saved/#view=csWVD3):

<img src="https://mintcdn.com/zylon/AAEt1MHLqi0nuOjf/images/operator-manual/hardware-requirements-images/basic.png?fit=max&auto=format&n=AAEt1MHLqi0nuOjf&q=85&s=7245534a38750a9db063c221ec94d7fa" alt="image.png" width="1854" height="1676" data-path="images/operator-manual/hardware-requirements-images/basic.png" />

This reference configuration should be adapted to meet the current GPU baseline for the **`baseline-96g` preset tier**. In practice, that means using a GPU such as the **NVIDIA RTX PRO 6000 (Workstation) (96 GB)**, an **A100 80GB**, an **H100**, or a multi-GPU setup that reaches a similar tier, together with a powerful CPU (16 cores), 128 GB of RAM, enough storage capacity to operate Zylon with margin to grow, a robust cooling solution, and a motherboard sized for the selected GPU layout.

Keep in mind that this is just a recommendation, so feel free to adapt it to your preferences while keeping similar capabilities for ideal performance. We have used Amazon as a provider considering that you can assemble all the parts together by yourself, but any provider that you usually work with should be able to get a similar hardware and assemble it for you.

## Reference hardware for big-size organization

In these scenarios, we don’t provide a reference hardware configuration until we understand the requirements not only regarding number of users, but also what kind of internal operations will be run in parallel by leveraging the platform API.

In you are in this situation, we are likely already discussing about this.
