remotely.living

Principal Server Engineer

UserWise Services · Remote - Hungary +4 locations · full-time · 2026-09-21

Apply for this job

Job description

About the Role

We're looking for a Principal Server Engineer to own the backend systems powering our live games - services currently supporting 1M+ DAU. You'll be responsible for performance, scalability, and reliability at the core of our infrastructure, working closely with the engineering team on everything from schema design to incident response.

Experience: 8+ years

Nice to Have

- CI/CD and DevOps experience (GitHub Actions, GitLab CI) plus monitoring/alerting (Prometheus, Datadog)

- Strong grasp of networking protocols — WebSockets, gRPC/HTTP2, UDP/TCP tuning for low-latency real-time systems

Requirements

Responsibilities

- Design, build, and maintain Go/Python backend services for live, high-traffic game titles, including server-authoritative logic and runtime modules on Nakama

- Extend and operate Nakama (Heroic Labs) for core game services: authentication, matchmaking, leaderboards, storage, RPCs, and realtime multiplayer

- Own database schema design and query performance across PostgreSQL systems, including Nakama's storage layer

- Drive horizontal scaling strategy (Nakama clustering, stateless services), optimizing CPU/memory footprint and reducing lock contention over vertical scaling

- Profile and resolve memory leaks, CPU spikes, and deadlocks in production

- Lead root-cause investigations for performance issues and outages, from detection through resolution

- Architect and maintain in-memory caching/state layers (Redis/KeyDB/Memcached)

- Design and run load tests against Nakama and supporting services to surface breaking points ahead of release

Requirements

Crucial

- Deep Go & Python expertise: language internals, execution models, and performance tuning (goroutine management, GC tuning, zero-allocation techniques)

- Strong PostgreSQL fluency: scalable schema design, query plan analysis, index strategy optimization, with real examples of resolving bottlenecks via data layout or query architecture changes (experience with Nakama's storage layer or CockroachDB a plus)

- Proven experience scaling systems supporting 350k+ DAU, with a horizontal-scaling-first philosophy (efficient serialization/deserialization via Protobuf/MsgPack, reduced lock contention)

- Hands-on profiling experience: identifying and resolving memory leaks, CPU spikes, and deadlocks

- At least 2 concrete examples of complex production performance issues or outages, including root cause analysis and resolution methodology

- At least 2 released production projects as a core backend engineer handling live traffic

Important

- Practical experience with Nakama (Heroic Labs): custom Go runtime modules (Lua/TypeScript a plus), matchmaker, storage engine, realtime sockets, and running Nakama clusters in production

- Hands-on experience with Redis, KeyDB, or Memcached: caching, rate-limiting, state management, including cache stampede mitigation, serialization overhead, and eviction policy design

- Solid TDD/BDD judgment and practical load-testing experience (e.g., Artillery), including realtime/socket workloads

- Practical GCP/Google App Engine experience: deploying, monitoring, and scaling serverless/container services

Benefits

- Market competitive, tax-free USD salaries

- Paid Time Off

- Annual Performance Reviews