About the Job

Join Impala AI, an innovative startup building a fully-managed, serverless LLM-inference platform, that enables data heavy enterprises to perform any AI task at any scale without limits.

Enterprises send us tokens; we handle everything behind the endpoint — model serving, GPU scheduling, autoscaling, batching, quantization, cost-per-token economics, and the SLAs that make inference safe to build a product on. No clusters to size, no GPUs to reserve, no capacity planning.

We're looking for our first Forward Deployed Engineer. You'll sit directly with our customers' engineering teams and get their workloads onto Impala — from the first scoping conversation to a production endpoint carrying real traffic, at a cost and latency profile they couldn't hit anywhere else.

This is an engineering role. You will write code every day, in customer repos and in ours. It also carries pieces of solutions architecture, product management, and pre-sales, and you should want that mix rather than tolerate it. Ambiguous business goals come in; observable, benchmarked, production services go out.

As the founding FDE you also define the function: what a POC looks like, what we promise and measure, which patterns get productized, and how the field feeds the roadmap. The next FDEs will work from what you build here.

What You'll Do

What You'll Bring