Skip to main content

Bjak

Backend Engineer, AI (Agent Systems)

United States (Remote)fulltimemidAdded 1 month ago

About this role

A1 is seeking a Backend Engineer to build and operate the inference and orchestration systems powering their AI assistant product. You'll own the critical layer between AI models and users, focusing on reliability, latency, and cost optimization for multi-step reasoning and real-world task completion.

What you'll do

  • Design and operate production inference pipelines and orchestration layers
  • Build stable, observable APIs serving AI features across mobile and desktop clients
  • Implement production monitoring, logging, alerting, and incident response systems
  • Optimize latency and throughput through caching, batching, and streaming strategies
  • Debug and resolve issues in distributed systems under load
  • Collaborate with ML and frontend systems to ensure seamless integration

What they're looking for

  • Backend engineering in production environments
  • High-throughput, low-latency service design
  • AI inference patterns (LLMs, embeddings, multimodal)
  • Python and/or Node.js
  • Kubernetes and Docker containerization
  • SQL and NoSQL databases
  • Distributed systems debugging
  • Production observability and incident response
Apply with Autofill

Opens the application — the Jobs AI extension fills it for you. Set up autofill

Opens the official application on the employer’s site. No login required.

Bjak

Bjak is a Southeast Asian fintech super app offering insurance, payments, savings, wallets, and investment products through a unified platform. The company is hiring full stack engineers, backend engineers, iOS developers, and Android engineers to build scalable features and reliable systems across mobile and web products.

Website
bjak.com
View all jobs at Bjak

Likely interview questions

  • Walk us through a time you designed an inference pipeline or orchestration layer for ML models in production. How did you handle latency and reliability concerns?
  • Describe your experience optimizing high-throughput, low-latency backend services. What specific bottlenecks did you identify and how did you resolve them?