jobgether logo

Staff Engineer (Core & MLOps)

jobgether

MexicoFULL_TIMEPosted 0 day(s) ago$0-$0 / yr

About jobgether

No company information provided.

About this Role.

This position is listed on behalf of a partner company, who manages all applications and next steps. Our partner is looking for a Staff Engineer (Core & MLOps) based in Mexico.

This role offers the opportunity to shape foundational infrastructure powering large-scale web data products and distributed engineering teams. You'll own the architecture of core control and context planes that enable services and AI-driven workflows to operate reliably and efficiently. Working across Kubernetes, Kafka, Java, Python, gRPC, and multi-cloud infrastructure, you'll tackle complex distributed systems challenges at production scale. You'll establish engineering standards, reliability practices, and service contracts that influence multiple product squads. The role combines hands-on architecture with technical leadership, mentoring, and cross-functional alignment. In a globally distributed, remote-first environment, you'll have significant autonomy to solve challenging infrastructure problems and influence long-term platform strategy.

Accountabilities:

  • Architect and evolve the control and context planes, advancing service and schema registries, SLO enforcement, health-aware routing, automated canary releases, and operational feedback loops.
  • Own the service chassis and golden path, maintaining and improving multi-language Java and Python client libraries, standardized workload specifications, Helm charts, and deployment pipelines.
  • Define and govern inter-service contracts, including gRPC and Protocol Buffer definitions, API gateway transcoding, versioning policies, and schema evolution standards.
  • Operate and improve the core platform infrastructure across Kubernetes, Terraform, HAProxy/Nginx, Confluent Kafka, real-time billing pipelines, Valkey, and database modernization initiatives.
  • Lead architectural strategy through Requests for Discussion (RFDs) covering workflow orchestration, gateway orchestration, multi-cluster routing, automated failover, and other critical platform initiatives.
  • Establish reliability engineering practices, including SLOs, SLIs, error budgets, fault isolation, and automated weighted canary deployments.
  • Participate in shared infrastructure on-call rotations, lead incident post-mortems, and convert operational insights into platform improvements.
  • Mentor engineers across multiple squads, review architectural proposals, and establish engineering practices that make reliable software development more consistent and efficient.
  • Requirements:
    • 10+ years of experience building scalable distributed backend systems, with a strong track record of creating internal platforms or core libraries adopted across engineering organizations.
    • Advanced Java expertise, including reactive frameworks such as Vert.x or Netty, combined with strong Python proficiency.
    • Deep experience with gRPC and Protocol Buffers, including schema evolution and backward compatibility in mission-critical systems.
    • Hands-on production experience with Kubernetes at scale, Terraform, and event-streaming platforms such as Kafka.
    • Experience designing automated telemetry pipelines, materialized views, feature stores, or other feedback systems that use production data to dynamically improve system behavior.
    • Strong reliability engineering background, including SLO/SLI definition, blast-radius analysis, fault tolerance, and rigorous service contracts.
    • Exceptional technical writing skills and the ability to communicate complex architectural concepts clearly while driving alignment across teams.
    • Strong written and interpersonal communication skills suited to a globally distributed, remote-first environment.
    • A curious, continuous-learning mindset with an interest in evaluating new technologies, architectures, and engineering approaches.
    • Experience with Temporal, DBOS, or similar durable execution platforms is a plus.
    • MLOps experience, including model serving, performance monitoring, or production drift detection, is advantageous.
    • Familiarity with zero-trust networking and service meshes such as SPIRE, mTLS, Cilium, Istio, or Envoy is beneficial.
    • Experience building developer tooling such as CLIs, SDKs, or project generators is a plus.
    • Experience with large-scale web scraping or crawling, or contributions to distributed-systems and data-extraction open-source projects, is advantageous.
    • Benefits:
      • Fully remote, remote-first working environment with flexible working hours.
      • Freedom and flexibility to work from the location where you are most productive.
      • Opportunity to work on core infrastructure supporting large-scale web data pipelines and distributed systems.
      • Exposure to cutting-edge open-source technologies, tools, and evolving AI and web data infrastructure.
      • Opportunities to attend conferences and connect with colleagues across the globe.
      • Collaboration with a diverse, multicultural, and globally distributed engineering community.
      • High level of autonomy and organizational trust.
      • Opportunities to influence platform architecture, engineering standards, and technical strategy across multiple teams.

Skills Required

Benefits & Perks