Senior AI Compute Infrastructure Engineer
Kraken · Panama
Descripcion del puesto
About the role
Kraken is building a dedicated AI Compute and Infrastructure team to power the next generation of model training, inference, evaluation, and experimentation across its exchange. As a Senior AI Compute Infrastructure Engineer you will own and operate GPU and accelerator clusters, ensuring they are fast, reliable, cost‑efficient and production‑grade.
Key responsibilities
- Own and operate GPU/accelerator clusters for training, inference, evaluation and experimentation, including drivers, runtimes, kernels, device plugins and node configuration.
- Design infrastructure that enables teams to run models locally on GPUs, reducing external dependency and controlling costs.
- Build and improve scheduling, orchestration, placement, quota management and utilization systems across heterogeneous accelerator environments.
- Optimize inference pipelines for latency, throughput, reliability and memory efficiency using frameworks such as vLLM, Triton Inference Server or TensorRT.
- Partner with ML engineers and researchers to remove bottlenecks in training, batch and online inference, deployment and debugging workflows.
- Develop observability for GPU utilization, memory pressure, queue depth, token throughput, request latency, failed workloads, capacity pressure and spend.
- Drive reliability through incident response, alerting, runbooks and post‑incident improvements for always‑on AI compute.
Required profile
- Senior‑level engineer with deep experience in GPU/accelerator infrastructure and cluster operations.
- Proven ability to work closely with AI/ML researchers, platform engineers, security and product teams.
- Strong focus on performance, reliability, cost discipline and production‑grade delivery.
Required skills
- GPU and accelerator cluster management
- Drivers, runtimes, kernels, device plugins
- Scheduling primitives, workload isolation, quota management
- vLLM, Triton Inference Server, TensorRT
- Observability metrics for GPU utilization, memory pressure, latency, throughput
- Capacity planning and cost‑efficient compute at scale
Questions fréquentes
Por que reporta esta oferta?
Explorar más
Salarios, guías y búsquedas en Panamá.
Postula en 30 segundos
Ingresa tu email para postular. Se creara una cuenta automaticamente.
Al continuar, aceptas nuestras condiciones de uso.
Ya tienes cuenta? Iniciar sesion
¿Una pregunta sobre esta oferta?
Hágala aquí: recibirá el resumen completo por correo, de inmediato.
Publicado hace 1 mes
Expira en 2 semanas
29 vistas · 0 interested
Aumenta tus posibilidades
Sube tu CV: te propondremos las ofertas que coinciden con tu perfil.
Analizando tu CV...
Kraken
Panama
Ofertas relacionadas
-
Digital Transformation Lead - Global Procurement
Nestle Operational Services Worldwide SA Panama -
Procurement Digital Innovation Project Manager
Nestle Operational Services Worldwide SA Panama -
Asistente en Programación de Informática
International Civil Aviation Organization Panama -
Asistente en Programación de Informática
UNDP Careers Panama -
Consultor DevSecOps Sr. – Proyecto 3 meses (Renovable)
Praxis Panama