Self-hosting an AI model means downloading its weights and running it on hardware you control, instead of sending requests to a provider's hosted API. The term applies most often to open-weight language models, but the same idea covers image, speech and embedding models used for document search.

It is closely related to on-premise AI, but "self-hosted" specifically describes the model itself — where it runs and who controls the file — while on-premises AI is the broader operational picture of running that model as part of a business system.


Where self-hosted models come from

Open-weight models are published by AI labs and research groups to public model repositories, where anyone can download the weights and run them on their own infrastructure. This is different from a proprietary hosted model, where only the provider can run it and everyone else accesses it through an API. Not every capable model is available this way — many of the strongest frontier models remain provider-hosted only — but a large and improving set of open-weight models is available to self-host.

Licensing is not one-size-fits-all

Every open-weight model ships with its own license, and licenses vary meaningfully: some permit unrestricted commercial use, others restrict use at scale, redistribution, or fine-tuning. "Open-weight" does not automatically mean "free to use however you like" — check the specific license of any model before deploying it in a business context, especially if you plan to fine-tune it or redistribute a modified version.

Storage and versioning, in practice

Model files range from a few gigabytes to well over a hundred, depending on size. Plan storage accordingly, and decide deliberately when to pin a specific model version versus upgrading — model publishers periodically release improved versions, and an unplanned automatic upgrade can change your system's behavior without warning if your deployment pulls the latest version by default.

What running one day to day requires

  • Enough compute to serve it at acceptable speed, sized to the model and your expected concurrent usage
  • A benchmarking process, since a model that tests well generally may not perform best on your specific documents or tasks
  • A monitoring and update plan, so you know when a model is underperforming and when a newer release is worth adopting
  • Access controls around who can query it and what data it can see, the same as any internal system handling business data

Do you need a data science team to self-host?

Not necessarily. Serving a well-chosen open-weight model for common tasks like document search or drafting assistance is largely an infrastructure and integration project, not a research project — you are not training a model from scratch. Where specialist skill matters most is in benchmarking candidates against your tasks and sizing the infrastructure correctly, which is exactly what a development partner typically handles.

The skills that matter most day to day, once a system is running, are closer to standard IT operations than machine learning research: monitoring a service, applying updates in a controlled way, and managing access — skills most internal IT teams already have or can pick up quickly with the right documentation from whoever built the system.

With a hosted API, the provider is responsible for keeping the model running, patched and available. Self-hosting shifts that responsibility to you, directly or through a support partner. That is not a reason to avoid it — it is simply a cost that needs to be planned for alongside the infrastructure itself, the same way any internal system carries an ongoing ownership cost beyond its initial build.

Where we fit

We select, benchmark and deploy self-hosted models for specific business tasks — sized to your hardware, licensed appropriately, and kept current with a periodic re-benchmarking process. See the on-premise AI overview for the full approach, or what makes one local model the best fit for a task for how we choose between candidates.

Frequently asked questions

Is self-hosting an AI model the same as on-premise AI?

They overlap but are not identical. Self-hosted describes running the model yourself rather than through a provider's API; on-premise AI is the broader operational picture of infrastructure, access control and lifecycle management around that model as part of a business system.

Do we need a data science team to self-host a model?

Not necessarily. You are not training a model from scratch — serving a well-chosen open-weight model is largely an infrastructure and integration task. Benchmarking candidates against your specific tasks is the part most worth specialist help.

Where do self-hosted AI models actually come from?

AI labs and research groups publish open-weight models to public repositories that anyone can download from and run on their own infrastructure, distinct from proprietary models that remain accessible only through the provider's own hosted API.

How large are the files for a self-hosted model?

It varies widely by model size, from a few gigabytes to well over a hundred. Plan storage capacity around the specific models you intend to run, and expect that keeping multiple model versions available takes meaningfully more space than one.

Do we need to fine-tune a model to self-host it usefully?

No. Many self-hosted deployments use a model as published, combined with retrieval over your documents, which is often enough for search, drafting and classification tasks without the added complexity of fine-tuning.