Infrastructure · GPU as a Service

Skip the hardware —reach GPU powerin minutes.

Bulutistan AI Cloud delivers L40S- and H200-class accelerators as a service for model training, fine-tuning, and high-volume inference. No procurement, setup, or maintenance overhead — pay only for what you use on an OpenStack-based sovereign cloud located in Türkiye.

L40S & H200 acceleratorsShared or dedicatedOpenStack-based sovereign cloudYour data stays in Türkiye
View LLM Service
141 GB
H200 HBM3e memory
4.8 TB/s
memory bandwidth
min
to provision
TR
data residency
Why GPU as a Service?

Don't buy the GPU — rent the outcome.

Building your own GPU cluster means long procurement lead times, heavy upfront investment, power-and-cooling planning, and constant maintenance. By the time the cards arrive, requirements have already shifted, and capacity ends up either idle or insufficient.

With GPU as a Service, you access accelerators as a service: it scales up in seconds when your workload grows and stops when it's done. Hardware CapEx turns into predictable OpEx that tracks your actual usage.

Your team stops managing clusters and focuses on the model, the product, and business value.

OpEx instead of CapEx

Instead of a million-dollar upfront hardware purchase, pay only for the GPU-hours you use as a predictable expense.

Up and running in minutes

No waiting on procurement or setup — spin up capacity instantly via a self-service API or console.

Scale to demand

Grow on the same infrastructure from PoC to high-volume production, and never pay for an idle card.

Zero hardware operations

Drivers, firmware, power, and cooling are on us — you focus on your workload.

GPU lineup

The right accelerator for the workload.

Not every workload wants the same GPU. Choose L40S for cost-effective inference and mid-scale training; choose H200 for memory-intensive large-model training and high-throughput inference. Both sit behind the same cloud and the same API.

NVIDIA L40S

Ada Lovelace · cost-effective inference & mid-scale training

Strong on price/performance. Ideal for production inference, mid-scale fine-tuning, computer vision, and image/video workloads; keeps unit cost low in shared scenarios.

Architecture
Ada Lovelace
GPU memory
48 GB GDDR6 (ECC)
Memory bandwidth
~864 GB/s
Tensor Core
4th gen · FP8
Typical workloads
Inference (7B–70B models)Mid-scale fine-tuningComputer vision & multimodalImage, video, and rendering

NVIDIA H200

Peak performance

Hopper · large-model training & memory-intensive inference

Breaks through the memory wall with 141 GB of HBM3e and 4.8 TB/s of bandwidth. Peak performance for LLM training, long-context and large-batch inference, and production traffic that demands high throughput.

Architecture
Hopper
GPU memory
141 GB HBM3e
Memory bandwidth
4.8 TB/s
Acceleration
Transformer Engine · FP8 · NVLink
Typical workloads
LLM trainingLong-context & large-batch inferenceMemory-intensive massive modelsHigh-volume production traffic
Deployment model

Shared flexibility or dedicated isolation.

Choose by the rhythm of your workload: a shared pool for variable and experimental loads, or single-tenant dedicated capacity for continuous, regulated production.

Shared GPU

Multi-tenant pool

On-demand capacity from an optimized pool. The most economical entry point for PoCs, variable traffic, and experimental workloads.

  • PAYG · $/1M token-based pricing
  • Capacity that scales in seconds
  • Ideal for PoCs and variable loads
  • Low cost of entry

Dedicated GPU

Single-tenant, isolated

Accelerators reserved and isolated for you. Predictable performance for continuous production, large-scale training, and enterprise scenarios that require a written SLA.

  • Full single-tenant isolation
  • Written SLA & reserved capacity
  • Stable for continuous production and training
  • KVKK / regulatory compliance add-ons
Sovereign cloud · OpenStack

Built on an open-standard cloud, with data sovereignty.

Bulutistan AI Cloud is built on open-source OpenStack. That means an infrastructure with no vendor lock-in, where your data stays within Türkiye's borders and capacity is under your control end to end.

GPU passthrough & MIG

GPUs on Nova compute are assigned to your workload as efficiently as possible — whether as full-card passthrough or partitioned with MIG.

Per-tenant isolation

Full isolation at the organization, project, and network level through the Keystone identity and project model; secure separation in a multi-tenant environment.

Self-service & IaC

Infrastructure as code (IaC) managed with Terraform and Kubernetes over the Nova, Neutron, and Cinder APIs.

Data sovereignty

Processing and storage stay in regions located in Türkiye; open standards preserve portability and auditability.

Infrastructure layers
05Application / model workload
04Kubernetes · Terraform (IaC)
03Nova · Neutron · Cinder · Keystone
02GPU passthrough · MIG partitioning
01L40S / H200 GPU pool

Every layer, from application to GPU, is built on open-standard components; it stays portable and auditable.

Workloads

One infrastructure for every compute-intensive load.

LLM training & fine-tuning

Train large language models from scratch or fine-tune them on your enterprise data; H200 memory makes long context and large batches possible.

High-volume inference

Serve production traffic at low latency; optimize unit cost with L40S and throughput with H200.

Computer vision & multimodal

Accelerated capacity for the training and inference workloads of image, video, and multimodal models.

RAG & vector workloads

GPU power that scales for embedding generation and high-volume RAG pipelines.

Research & PoC

From idea to prototype; fast, cheap, and reproducible experimentation environments on the shared pool.

HPC & simulation

Run parallel workloads like scientific computing, simulation, and rendering at GPU speed.

Why Bulutistan

From GPU to model, under one roof.

Infrastructure (GPU) and LLM Service live on the same cloud: a single end-to-end provider from hardware to production-ready model output.

End-to-end AI Cloud

GPU as a Service and LLMaaS under one roof; a frictionless path from infrastructure to model API.

Data sovereignty in Türkiye

Processing and storage in local regions; a foundation aligned with KVKK and enterprise regulatory expectations.

Usage-based economics

Predictable cost proportional to usage instead of an upfront hardware investment; no paying for idle capacity.

Enterprise isolation & support

Go to production with confidence, backed by project-level isolation, dedicated capacity, and solution engineer support.

Request a call

Let's build the right GPU plan for your workload together.

Share your capacity needs, model size, and budget; our solutions team will work with you to determine whether shared or dedicated, L40S or H200 is the right fit. Pricing and quotas are confirmed in writing during the call.