Skip to content

On-premise AI hardware

AI server in your server room

Sending company documents to an external API can be unacceptable legally or simply out of caution. We deploy an on-premises machine with 128 GB GPU memory running open-source language models with no cloud connection.

In practice

Data stays in the building

An on-premise LLM server is a single unit in a rack cabinet. It delivers the same performance as a cloud service, but without sending any data outside.

A configuration with 128 GB of VRAM runs quantized 70B class models today. That is usually enough for a dozen concurrent conversations in a mid-sized company.

We start with a single department, most often sales proposals or technical documentation. We measure the actual time saved before connecting additional teams. After a two-day training, your IT specialist adds models and creates accounts independently.

The hardware belongs to you from day one. If you change software providers after a year, the machine stays in place.

  • How it works

    The machine arrives at your server room ready to work: operating system, drivers, and several open-source models already installed. Prompts and completions are stored on your local drive, and the default configuration blocks all outbound traffic.

  • Integrations

    The server exposes a standard API, so we connect it to the tools you already use: email, CRM, network drive, spreadsheets. We mirror document permissions from your systems so no user sees files in responses they lack access to.

  • Who decides

    The admin password stays with you, and your company remains the data controller. We only access the server when someone on your team grants us entry. The audit log tracks every query, and the SLA defines who answers the phone if hardware fails.

Problem

Concerned about sending company documents to the cloud?

HEXART solution

A server with 128 GB of VRAM sits on your premises and runs quantized 70B class models. Queries remain on your local drive, and access is managed by your administrator. We start with a single process and measure time saved before expanding the deployment.

Ask about server configuration

Questions and answers

Frequently asked questions

What models fit on a server with 128 GB of VRAM?

Quantized 70B class models, which is usually enough for a dozen concurrent conversations in a mid-sized company.

Do queries leave the company?

No. Prompts and completions are stored on the local drive, and the default configuration blocks all outbound traffic. The admin password stays with you, and your company remains the data controller.

Do we own the hardware?

Yes, from day one. If you change software providers after a year, the machine stays in place. The SLA defines who answers the phone if hardware fails.

Who manages the server after deployment?

Your IT specialist. After a two-day training, they add models and create accounts independently. The server exposes a standard API, so we connect it to the email, CRM, network drive, and spreadsheets you already use.

No-obligation call

Schedule a workshop with Paul

Thirty minutes or a full workshop, choose what suits you. Book a slot directly in the calendar and get instant confirmation.

Paul Lazniak

HEXART Founder

Contact

Let us talk about your project

We will show you a working product before we start talking about it. Let us know what you need: a film, an XR environment, an AI system, or brand identity.

Write or call

A few sentences are enough: what needs to be created, by when and for whom.

Address
Aleja Zwycięstwa 96/98, 81-451 Gdynia

Deployments catalog

View full catalogue
  • On-premise server knowledge base (RAG)

    Search contracts, technical documentation, and correspondence via chat when company policy prohibits sending them to the cloud. The model and index run in your rack cabinet. A single machine handles 5 to 20 concurrent simple queries within a 100 to 1,000 ms window. An additional server doubles that capacity.

  • Encrypted internal messenger with AI

    Acquisition plans, salary negotiations, and HR matters circulate today on messengers whose terms of service no one has read. We deploy a dedicated channel: a server on your premises, end-to-end encryption, and an AI assistant on the same network.

  • AI agents for chat and phone

    A large share of support tickets are repeated questions about order status, invoices, and delivery times. An AI agent answers them directly from your documentation, resolving 30–60% of such inquiries in our deployments. It routes the rest to a human agent along with full conversation context.

  • AI contract analysis before signing

    A 40-page contract usually contains 5 clauses that pose real risks. The system locates them and displays paragraph numbers so the lawyer or responsible person can focus on what matters first.

  • Company social media automation

    A company profile usually goes quiet not from a lack of ideas, but because no one has an hour for it on a Wednesday. The system drafts a weekly batch of posts with visuals, and you approve, edit, or reject them.