Open-source LLM hosting
Run Llama, Mistral, Gemma, Qwen and similar models on hardware you control.
Models
Llama family, Mistral, Gemma, fine-tuned variants
GPU / AI Compute
GPU compute for open-source LLMs, computer vision, fine-tuning, and embeddings. Live pricing below. GPU stock is allocated per order and provisioned by hand, so setup is not instant.
GPU Compute
Live pricing, order today
NVIDIA
RTX-class GPUs
4 tiers
Spark to Reactor
Per-order
Provisioned by hand
API
Provisioning + monitoring
The full ladder
Forge
NVIDIA RTX PRO 6000 · 96 GB VRAM · 256 GB RAM
Managed /mo: $6,820.00
Foundry
NVIDIA RTX PRO 6000 · 96 GB VRAM · 512 GB RAM
Managed /mo: $10,220.00
Reactor
NVIDIA RTX PRO 6000 · 96 GB VRAM · 768 GB RAM
Managed /mo: $13,140.00
Limited, reserve when short
What you'll run on it
These are the workloads the GPU tiers are built for. If yours matches, pick the tier that fits and order it.
Run Llama, Mistral, Gemma, Qwen and similar models on hardware you control.
Models
Llama family, Mistral, Gemma, fine-tuned variants
Image classification, object detection, segmentation. Training and inference on the same platform without re-deploying.
Models
YOLO, SAM, custom PyTorch / TensorFlow models
LoRA / QLoRA fine-tunes on open-base models. Per-hour billing means you only pay for actual training time.
Models
LoRA, QLoRA, full fine-tunes on smaller models
Generate embeddings at scale for retrieval-augmented generation pipelines. Connect to your vector DB.
Models
sentence-transformers, BGE, e5, custom embedders
Stable Diffusion, FLUX, Kandinsky. Batch-generate or expose as an API endpoint.
Models
SD, SDXL, FLUX, ControlNet, custom checkpoints
Speech-to-text, text-to-speech, voice cloning. Whisper transcription pipelines.
Models
Whisper, Coqui, Bark, custom TTS
Available right now
Our Cloud VPS lineup runs JupyterHub as a one-click installer. CPU-bound inference, smaller models, embedding generation, batch pipelines: all run today without waiting for GPU.
FAQ
Order any tier today at the price shown. GPU stock is allocated per order and provisioned by hand (Robot-class hardware), so there is a short lead time between order and handover rather than instant self-provisioning. We confirm timing at checkout.
NVIDIA GPUs across four tiers: Spark on an RTX 4000, and Forge, Foundry, and Reactor on the RTX PRO 6000 with increasing system memory. Each tier lists its GPU, VRAM, and RAM on the pricing grid above.
We confirm the exact data-centre location with you before handover. GPU stock is allocated per order and provisioned by hand, so we are transparent about the region and specifics at launch.
Live monthly pricing is on this page, per tier, in your currency. Both unmanaged and managed prices are shown.
Yes. Spin up a Cloud VPS with JupyterHub one-click installed. CPU-bound LLM inference (smaller models), embedding generation, batch pipelines, all run there today. When GPU lands you migrate the code, no rewrite.
Training large foundation models from scratch (we are targeting fine-tuning and inference, not the >$10M training-run scale). Closed-source proprietary models requiring vendor agreements. Anything requiring InfiniBand-scale multi-GPU clusters at launch (multi-GPU is roadmap, not GA).
Stock is allocated per order and provisioned by hand, so there can be a short lead time between order and handover. This is the same manual-fulfilment path as our dedicated and storage server lines.
Kora knows the GPU compute product: the hardware, the pricing, the supported frameworks, and the deploy patterns.
GPU / AI Compute
Live monthly pricing per tier, root access. Stock is allocated per order and provisioned by hand, so setup is not instant.