Senior ML Compiler Engineer
Job Description:
Para nosotros es fundamental que sean Compiler Engineers reales. El principal filtro técnico es experiencia hands-on con MLIR y/o LLVM, acompañada de experiencia en optimizing compilers, C++ y optimización para algún hardware target (GPU, TPU, FPGA, CGRA, DSP, AI ASIC, etc.).Nos interesa especialmente encontrar evidencia de trabajo con IRs, lowering, compiler/graph transformations, optimization passes, scheduling, memory hierarchy y data movement.Importante: perfiles principalmente de Data Science, Applied ML, MLOps, Backend/C++, FPGA/RTL o CUDA/kernel development no serían suficientes si no tienen experiencia concreta en compiler engineering.
We are looking for a Senior ML Compiler Engineer to join a specialized engineering team working on next-generation AI
accelerators for a US-based AI infrastructure company.
Unlike a GPU, a reconfigurable dataflow architecture does not execute a fixed instruction set. An entire model graph is
spatially mapped onto the chip’s compute and memory fabric, and the compiler is what makes that possible, and what
makes it fast. It takes a model authored in a standard framework and, without asking the user to hand-tune kernels,
extracts the dataflow graph, fuses and pipelines operators, allocates on-chip memory, and maps computation onto the
hardware.
This is not GPU kernel work and it is not applied ML. It is production compiler engineering on custom silicon, at
datacenter scale. You will work at the layer where inference cost, latency, and energy efficiency are actually decided.
This is a fully remote role from Latin America, working US-aligned hours. No relocation required.
��️ Your Impact and Responsibilities
• Design, build, and optimize passes in an ML compiler stack that lowers model graphs from PyTorch or ONNX front
ends into an optimized, hardware-specific dataflow representation.
• Develop graph-level transformations, including operator fusion, tiling, meta-pipelining, parallelization, and
memory-aware scheduling, that maximize on-chip data reuse and minimize data movement.
• Build cost models and heuristics so that fusion, tiling, and scheduling decisions are made automatically, preserving
push-button compilation for end users.
• Diagnose bottlenecks across the compiler, runtime, and hardware boundary, and drive the fix wherever it lives: a
compiler pass, the IR, or the runtime.
• Contribute to compiler infrastructure quality: IR design, testing, numerical correctness, and performance
regression tracking.
• Collaborate with hardware, kernel, runtime, and ML research teams to co-design compiler features against
evolving silicon and model workloads.
✅ What We’re Looking For
• 5+ years of production compiler or high-performance systems engineering, with meaningful work on an
optimizing compiler (front-end, middle-end, or back-end).
• Hands-on experience with MLIR and/or LLVM, or a comparable domain-specific compiler framework with real
depth.
• Strong compiler fundamentals: intermediate representations, dataflow analysis, graph transformations, loop and
operator optimization, scheduling, and memory allocation.
• Modern C++ at a production level, plus Python.
• A working understanding of deep learning fundamentals and how transformer models lower to hardware.
• Experience optimizing for at least one hardware target (GPU, TPU, FPGA, CGRA, DSP, or a custom AI ASIC),
including reasoning about memory hierarchy, parallelism, and data movement.
• A measurable track record of making workloads meaningfully faster, and the profiling discipline to prove it.
• Advanced English (C1). Daily collaboration with a US engineering team.
⭐ Nice to Have
• Experience with a dataflow, spatial, or coarse-grained reconfigurable architecture (CGRA), including its place-and-
route problem.
• Experience with ML compiler stacks such as XLA, TVM, IREE, Torch-MLIR, torch.fx, Triton, Glow, or a proprietary
equivalent.
• Familiarity with operator fusion strategies, kernel scheduling, and quantization or mixed-precision execution.
• Experience bringing up large language models, including sharding, pipelining, and multi-chip execution.
• Open-source compiler contributions (LLVM, MLIR, IREE, TVM, Triton) or publications in compilers, computer
architecture, or ML systems.
A note on what we do not require: a PhD, or prior experience with this specific accelerator. If you have real compiler
depth, we want to talk to you.
�� Why Join Applica?
Because we are building the future of technology alongside people who want to create impact, grow professionally, and
be part of a team committed to innovation, continuous learning, and excellence.
Here you will have the opportunity to work on genuinely hard problems at the frontier of AI infrastructure, collaborate
with top-level professionals, and grow your career in a dynamic, flexible environment with global reach.
✨ What We Offer
�� First-class tools and technology to do your work.
�� Training, certifications, and continuous learning opportunities.
�� A flexible, results-oriented way of working.
��️ Paid time off so you can balance your personal and professional life.
�� Real opportunities for growth and career development.
�� A collaborative, inclusive, people-centered culture.
�� Referral program.
�� Celebrations, recognition, and benefits on special occasions.