Senior ML Compiler Engineer

  • Buenos Aires, Argentina
  • Full-Time
  • Remote

Job Description:

Para nosotros es fundamental que sean Compiler Engineers reales. El principal filtro técnico es experiencia hands-on con MLIR y/o LLVM, acompañada de experiencia en optimizing compilers, C++ y optimización para algún hardware target (GPU, TPU, FPGA, CGRA, DSP, AI ASIC, etc.).Nos interesa especialmente encontrar evidencia de trabajo con IRs, lowering, compiler/graph transformations, optimization passes, scheduling, memory hierarchy y data movement.Importante: perfiles principalmente de Data Science, Applied ML, MLOps, Backend/C++, FPGA/RTL o CUDA/kernel development no serían suficientes si no tienen experiencia concreta en compiler engineering.


We are looking for a Senior ML Compiler Engineer to join a specialized engineering team working on next-generation AI

accelerators for a US-based AI infrastructure company.

Unlike a GPU, a reconfigurable dataflow architecture does not execute a fixed instruction set. An entire model graph is

spatially mapped onto the chip’s compute and memory fabric, and the compiler is what makes that possible, and what

makes it fast. It takes a model authored in a standard framework and, without asking the user to hand-tune kernels,

extracts the dataflow graph, fuses and pipelines operators, allocates on-chip memory, and maps computation onto the

hardware.

This is not GPU kernel work and it is not applied ML. It is production compiler engineering on custom silicon, at

datacenter scale. You will work at the layer where inference cost, latency, and energy efficiency are actually decided.

This is a fully remote role from Latin America, working US-aligned hours. No relocation required.

��️ Your Impact and Responsibilities

• Design, build, and optimize passes in an ML compiler stack that lowers model graphs from PyTorch or ONNX front

ends into an optimized, hardware-specific dataflow representation.

• Develop graph-level transformations, including operator fusion, tiling, meta-pipelining, parallelization, and

memory-aware scheduling, that maximize on-chip data reuse and minimize data movement.

• Build cost models and heuristics so that fusion, tiling, and scheduling decisions are made automatically, preserving

push-button compilation for end users.

• Diagnose bottlenecks across the compiler, runtime, and hardware boundary, and drive the fix wherever it lives: a

compiler pass, the IR, or the runtime.

• Contribute to compiler infrastructure quality: IR design, testing, numerical correctness, and performance

regression tracking.

• Collaborate with hardware, kernel, runtime, and ML research teams to co-design compiler features against

evolving silicon and model workloads.

✅ What We’re Looking For

• 5+ years of production compiler or high-performance systems engineering, with meaningful work on an

optimizing compiler (front-end, middle-end, or back-end).

• Hands-on experience with MLIR and/or LLVM, or a comparable domain-specific compiler framework with real

depth.

• Strong compiler fundamentals: intermediate representations, dataflow analysis, graph transformations, loop and

operator optimization, scheduling, and memory allocation.

• Modern C++ at a production level, plus Python.

• A working understanding of deep learning fundamentals and how transformer models lower to hardware.

• Experience optimizing for at least one hardware target (GPU, TPU, FPGA, CGRA, DSP, or a custom AI ASIC),

including reasoning about memory hierarchy, parallelism, and data movement.

• A measurable track record of making workloads meaningfully faster, and the profiling discipline to prove it.

• Advanced English (C1). Daily collaboration with a US engineering team.

⭐ Nice to Have

• Experience with a dataflow, spatial, or coarse-grained reconfigurable architecture (CGRA), including its place-and-

route problem.

• Experience with ML compiler stacks such as XLA, TVM, IREE, Torch-MLIR, torch.fx, Triton, Glow, or a proprietary

equivalent.

• Familiarity with operator fusion strategies, kernel scheduling, and quantization or mixed-precision execution.

• Experience bringing up large language models, including sharding, pipelining, and multi-chip execution.

• Open-source compiler contributions (LLVM, MLIR, IREE, TVM, Triton) or publications in compilers, computer

architecture, or ML systems.

A note on what we do not require: a PhD, or prior experience with this specific accelerator. If you have real compiler

depth, we want to talk to you.

�� Why Join Applica?

Because we are building the future of technology alongside people who want to create impact, grow professionally, and

be part of a team committed to innovation, continuous learning, and excellence.

Here you will have the opportunity to work on genuinely hard problems at the frontier of AI infrastructure, collaborate

with top-level professionals, and grow your career in a dynamic, flexible environment with global reach.

✨ What We Offer

�� First-class tools and technology to do your work.

�� Training, certifications, and continuous learning opportunities.

�� A flexible, results-oriented way of working.

��️ Paid time off so you can balance your personal and professional life.

�� Real opportunities for growth and career development.

�� A collaborative, inclusive, people-centered culture.

�� Referral program.

�� Celebrations, recognition, and benefits on special occasions.