
Fast-track your ML job hunt :
As a Research Engineer on Forge, you will turn real customer requirements into reliable training and deployment workflows. You’ll work end‑to‑end across model adaptation and post‑training (CPT/SFT/RL/distillation), evaluation, data, and infrastructure. The role bridges research experimentation and production constraints.
This role sits in Applied Science, with direct impact on client outcomes. You’ll collaborate closely with scientists, engineers, product, and customer‑facing teams to ensure Forge projects ship, are maintainable, and can be trusted by others.
Interview focus can vary (algorithms, infrastructure, evals, or data). You don’t need to match every bullet below to apply.
Build and improve post‑training and evaluation workflows (CPT/SFT/RL/distillation), turning prototypes into repeatable Forge “recipes”.
Develop tools and pipelines for synthetic data generation, data curation, training, evaluation, and deployment.
Debug and harden large‑scale ML systems: distributed training, scheduling/execution, checkpointing, observability, and reproducibility.
Improve the Forge codebase via clear APIs, tests, documentation, and maintainable abstractions.
Push the frontier of our RL training stack (e.g., high-throughput async rollout and scalable post‑training systems at frontier-model scale)
Make sure Forge deployment is seamless and adaptable to a diversity of clients (hardware access, software stack, cloud and on-premises, …)
Partner with researchers and infrastructure engineers to translate bottlenecks into concrete system improvements.
Strong Python engineering skills and experience working in large codebases (testing, code review, CI, operational ownership).
Hands‑on experience with PyTorch, JAX, or similar.
Strong systems and infrastructure fundamentals.
Experience with LLM training or post‑training: fine‑tuning, RL, distillation, evaluation, and/or data pipelines.
Excellent debugging skills in ambiguous systems (distributed jobs, data issues, quality regressions, infra failures).
Clear communication with technical and non‑technical stakeholders.
High agency, low ego, and comfort in fast‑moving, under‑specified environments.
Distributed training experience (FSDP, DeepSpeed, Megatron, etc.).
Cluster/orchestration experience (SLURM, Ray, Kubernetes, Kueue, Karpenter, Skypilot, etc.).
Experience building reliable ML infrastructure, evaluation systems, or large‑scale data processing pipelines.
Research experience in LLMs, agents, multimodal models, reasoning, code, or domain adaptation.
Open‑source contributions, publications, or widely used internal tooling.
Experience training multi‑billion‑parameter models (pre‑training or RL). Experience training on petabyte- and exabyte-scale datasets.
Ability to identify bottlenecks across the stack and drive improvements from first principles.
Fast-track your ML job hunt :