remotely.living

GPU & ML Infrastructure Engineer

Svitla Systems · Remote - Argentina, Chile, Mexico, United States · 2026-08-06

Apply for this job

Job description

Svitla Systems Inc. is looking for a GPU & ML Infrastructure Engineer for a full-time position (40 hours per week) in the USA. Our client is a stealth startup. The successful candidate will own the end-to-end data generation process for GPU systems, including benchmark workloads, automated deployment, hardware telemetry collection, data quality validation, and dataset delivery. The role requires working with multiple NVIDIA data-center GPU generations and embedded or edge platforms.

Requirements

- Strong experience deploying LLM inference and training workloads on GPUs, including quantized models.

- Experience diagnosing sensor, logging, and sampling issues in time-series hardware data.

- Ability to build reproducible GPU workloads and control sources of run-to-run variation.

- Strong Linux systems knowledge, including GPU driver stacks, process orchestration, scheduling, and timing.

- Experience collecting hardware telemetry programmatically using NVML, DCGM, BMC, IPMI, or Redfish.

- Strong Python skills for automation, telemetry collection, and data processing.

- Experience automating workload deployment and data collection across different hardware platforms.

- Ability to work independently and take ownership of technical processes.

- Availability to overlap with the client’s team until 11:00 a.m. PST.

Nice to have

- Experience building data collection pipelines for hardware testing or systems research.

- Familiarity with GPU benchmarking, stress testing, and benchmark methodology.

- Knowledge of GPU power, thermal management, multi-GPU scaling, and NCCL.

- Experience building automated data quality checks for time-series or sensor data.

Responsibilities

- Port existing test procedures to new data-center GPUs and edge devices.

- Build and maintain GPU benchmark workloads, including synthetic kernels and LLM inference and training.

- Ensure telemetry collection is accurate and consistent across platforms and data sources.

- Investigate sampling issues, timestamp inconsistencies, missing sensor data, and logging anomalies.

- Automate workload deployment, execution, data collection, and environment cleanup.

- Build automated checks for missing samples, irregular intervals, clock mismatches, and invalid telemetry.

- Deliver datasets in a consistent and documented format with complete run metadata.

- Document hardware targets, configurations, procedures, driver versions, and firmware versions.