About the Job
Join Impala AI, an innovative startup building a fully-managed, serverless LLM-inference platform, that enables data heavy enterprises to perform any AI task at any scale without limits.
Enterprises send us tokens; we handle everything behind the endpoint — model serving, GPU scheduling, autoscaling, batching, quantization, cost-per-token economics, and the SLAs that make inference safe to build a product on. No clusters to size, no GPUs to reserve, no capacity planning.
We're looking for our first Forward Deployed Engineer. You'll sit directly with our customers' engineering teams and get their workloads onto Impala — from the first scoping conversation to a production endpoint carrying real traffic, at a cost and latency profile they couldn't hit anywhere else.
This is an engineering role. You will write code every day, in customer repos and in ours. It also carries pieces of solutions architecture, product management, and pre-sales, and you should want that mix rather than tolerate it. Ambiguous business goals come in; observable, benchmarked, production services go out.
As the founding FDE you also define the function: what a POC looks like, what we promise and measure, which patterns get productized, and how the field feeds the roadmap. The next FDEs will work from what you build here.
What You'll Do
- Own customer outcomes end to end: problem framing, evaluation design, migration, deployment, benchmarking, monitoring. You're the technical owner from first call through expansion.
- Win the POC: turn a vague objective into a tight spec and a working proof of concept fast, with explicit quality, latency, throughput, and cost-per-token targets — and hit them.
- Tune serverless inference for real workloads: model and engine selection, batching strategy, KV cache behavior, speculative decoding, quantization, parallelism, cold-start and autoscaling behavior under bursty traffic. Diagnose regressions down to the inference engine.
- Migrate workloads onto Impala: move customers off OpenAI-compatible APIs, self-managed vLLM, SageMaker, or their own GPU fleets — and prove out the quality and cost delta with numbers.
- Guide model strategy: advise on open-weight model selection, distillation, and fine-tuning for specific tasks; help customers get from a general-purpose frontier model to a smaller, faster, cheaper one that holds quality.
- Close the product loop: bring the field back into the roadmap — write the PRDs, land the PRs, and turn one-off customer work into platform features.
- Be the technical anchor in the room: support sales on complex evaluations, run technical onboarding, earn trust with staff engineers and CTOs.
What You'll Bring
- 4+ years building and shipping production software
- Inference in production: hands-on experience serving LLMs with vLLM, SGLang, TensorRT-LLM or equivalent, and real intuition for what makes inference fast or expensive.
- Optimization fundamentals: working knowledge of batching, KV cache, quantization, speculative decoding, tensor and pipeline parallelism — and the tradeoffs between them.
- Model judgment: fluency with the open-weight model landscape and good instincts on model selection for a given task, hardware profile, and latency budget.
- Production cloud comfort: containers, Kubernetes, observability, CI/CD. You don't need to be an SRE, but nothing here should be a black box.