Jobiglo

No results.

Senior AI Compute Infrastructure Engineer

Kraken · Panama

Senior 🇬🇧 English
Drivers Runtimes Kernels Device plugins Scheduling primitives Workload isolation Quota management VLLM Triton Inference Server TensorRT Capacity planning

Job description

About the role

Kraken is building a dedicated AI Compute and Infrastructure team to power the next generation of model training, inference, evaluation, and experimentation across its exchange. As a Senior AI Compute Infrastructure Engineer you will own and operate GPU and accelerator clusters, ensuring they are fast, reliable, cost‑efficient and production‑grade.

Key responsibilities

  • Own and operate GPU/accelerator clusters for training, inference, evaluation and experimentation, including drivers, runtimes, kernels, device plugins and node configuration.
  • Design infrastructure that enables teams to run models locally on GPUs, reducing external dependency and controlling costs.
  • Build and improve scheduling, orchestration, placement, quota management and utilization systems across heterogeneous accelerator environments.
  • Optimize inference pipelines for latency, throughput, reliability and memory efficiency using frameworks such as vLLM, Triton Inference Server or TensorRT.
  • Partner with ML engineers and researchers to remove bottlenecks in training, batch and online inference, deployment and debugging workflows.
  • Develop observability for GPU utilization, memory pressure, queue depth, token throughput, request latency, failed workloads, capacity pressure and spend.
  • Drive reliability through incident response, alerting, runbooks and post‑incident improvements for always‑on AI compute.

Required profile

  • Senior‑level engineer with deep experience in GPU/accelerator infrastructure and cluster operations.
  • Proven ability to work closely with AI/ML researchers, platform engineers, security and product teams.
  • Strong focus on performance, reliability, cost discipline and production‑grade delivery.

Required skills

  • GPU and accelerator cluster management
  • Drivers, runtimes, kernels, device plugins
  • Scheduling primitives, workload isolation, quota management
  • vLLM, Triton Inference Server, TensorRT
  • Observability metrics for GPU utilization, memory pressure, latency, throughput
  • Capacity planning and cost‑efficient compute at scale

Questions fréquentes

Le salaire n'est pas communiqué publiquement par le recruteur. Vous pouvez postuler et négocier directement avec Kraken.
Cliquez sur "Postuler maintenant" en haut de la page. Vous pouvez importer votre CV en 1 clic — Jobiglo extrait automatiquement vos informations et postule pour vous.

Why are you reporting this job?

Thank you for your report. We will review this job.

Apply in 30 seconds

Enter your email to apply. An account will be created automatically.

By continuing, you accept our terms of use.

Already have an account? Login

A question about this job?

Ask it here: you will get the full job summary by e-mail, right away.

💬 Chat with us on Telegram

Published 1 month ago

Expires 2 weeks from now

28 views · 0 interested

Boost your chances

Upload your CV — we will match you with relevant openings.

Analyzing your CV...

Kraken

Panama