GPU & ML Infrastructure Engineer
Svitla Systems · Remote - Argentina, Chile, Mexico, United States · 2026-08-06
Job description
Svitla Systems Inc. is looking for a GPU & ML Infrastructure Engineer for a full-time position (40 hours per week) in the USA. Our client is a stealth startup. The successful candidate will own the end-to-end data generation process for GPU systems, including benchmark workloads, automated deployment, hardware telemetry collection, data quality validation, and dataset delivery. The role requires working with multiple NVIDIA data-center GPU generations and embedded or edge platforms.
Requirements
- Strong experience deploying LLM inference and training workloads on GPUs, including quantized models.
- Experience diagnosing sensor, logging, and sampling issues in time-series hardware data.
- Ability to build reproducible GPU workloads and control sources of run-to-run variation.
- Strong Linux systems knowledge, including GPU driver stacks, process orchestration, scheduling, and timing.
- Experience collecting hardware telemetry programmatically using NVML, DCGM, BMC, IPMI, or Redfish.
- Strong Python skills for automation, telemetry collection, and data processing.
- Experience automating workload deployment and data collection across different hardware platforms.
- Ability to work independently and take ownership of technical processes.
- Availability to overlap with the client’s team until 11:00 a.m. PST.
Nice to have
- Experience building data collection pipelines for hardware testing or systems research.
- Familiarity with GPU benchmarking, stress testing, and benchmark methodology.
- Knowledge of GPU power, thermal management, multi-GPU scaling, and NCCL.
- Experience building automated data quality checks for time-series or sensor data.
Responsibilities
- Port existing test procedures to new data-center GPUs and edge devices.
- Build and maintain GPU benchmark workloads, including synthetic kernels and LLM inference and training.
- Ensure telemetry collection is accurate and consistent across platforms and data sources.
- Investigate sampling issues, timestamp inconsistencies, missing sensor data, and logging anomalies.
- Automate workload deployment, execution, data collection, and environment cleanup.
- Build automated checks for missing samples, irregular intervals, clock mismatches, and invalid telemetry.
- Deliver datasets in a consistent and documented format with complete run metadata.
- Document hardware targets, configurations, procedures, driver versions, and firmware versions.