Model Routing & Inference Team Lead

c220dc7b-fb8 Model Routing & Inference Team Lead About the Role

You will lead the Model Routing & Inference team at Cursor, owning the inference platform that powers every AI interaction in the product. This team owns the full inference path: making Cursor's AI faster, more reliable, and more cost-effective at a scale few teams in the world get to operate at. Every agent session, every tab completion, and every chat message flows through your stack.

You'll set technical direction for cluster management, inference optimization, and traffic egress, building the platform that lets the rest of the company move fast without worrying about provider complexity. You'll lead a team of strong engineers, set strong direction for the business, and make the calls that balance latency, cost, reliability, and user experience across millions of daily requests.

What you’ll do

Building and evolving our inference gateway, a single abstraction over every provider's API semantics, so model onboarding becomes a config change.
Building the systems that dynamically select the best model for each request based on cost, latency, and quality.
Managing GPU cluster utilization and capacity planning across providers, optimizing for cost and performance.
Designing routing backpressure and admission control so traffic spikes don't cascade into providers.
Hiring and growing the team: sourcing, interviewing, and closing top inference and systems talent, while developing your engineers through coaching, mentorship, and high-leverage project assignments.

You may be a fit if

You have led engineering teams building high-throughput, low-latency distributed systems, especially in inference serving, traffic routing, or real-time data pipelines.
You're comfortable reasoning about cost/performance tradeoffs at scale (GPU utilization, provider economics, capacity planning) and making decisions with incomplete information.
You have strong software engineering fundamentals and enjoy shipping production systems that handle millions of requests.
Experience with model serving frameworks (vLLM, TensorRT-LLM, TGI), load balancing, or building resilient multi-provider architectures is a plus.
You make good calls in the gray area: weighing reliability, cost, latency, and user experience when there isn't a single 'right' answer.

XML job scraping automation by YubHub

]]> full-time senior remote model serving frameworks, load balancing, resilient multi-provider architectures, cost/performance tradeoffs, GPU utilization Engineering Technology Cursor https://logos.yubhub.co/cursor.com.png Cursor is an AI-powered product company that operates at a large scale. https://cursor.com https://cursor.com/careers/engineering-manager-model-routing-inference 2026-04-24 96d93009-885 ChatGPT Performance Engineer OpenAI is looking for an experienced Performance Engineer to help us scale the performance, reliability, and efficiency of our systems. In this role, you'll apply deep technical expertise to optimize infrastructure and application-level performance across mission-critical products like ChatGPT and our developer API.

We are looking for engineers who thrive in ambiguous environments, value deep systems understanding, and are motivated by delivering measurable impact. This is a highly technical, individual contributor role focused on root-cause analysis, profiling, instrumentation, and architecture-level performance improvements across our stack.

Key responsibilities include:

Analyzing and optimizing performance across application, middleware, runtime, and infrastructure layers,networking, storage, Python runtime, GPU utilization, and beyond.
Developing tooling and metrics that provide deep observability into system performance.
Collaborating closely with infra, platform, training, and product teams to identify key performance goals and drive systemic improvements.
Influencing architecture and design decisions to prioritize latency, throughput, and efficiency at scale.
Leading investigations into high-impact performance regressions or scalability issues in production.
Driving performance testing strategies and helping define SLAs/SLOs around latency and throughput for critical systems.

If you have 7+ years of experience in software engineering with a strong track record in performance or reliability of high-scale distributed systems, you might thrive in this role. You should be deeply comfortable with performance profiling tools and tracing systems, have experience optimizing performance across one or more layers of the stack, and have a strong understanding of OS internals, scheduling, memory management, and IO patterns.

XML job scraping automation by YubHub

]]> Full time senior remote $325K – $405K performance engineering, distributed systems, python, gpu utilization, os internals, scheduling, memory management, io patterns Engineering Technology OpenAI https://logos.yubhub.co/openai.com.png OpenAI is an AI research and deployment company dedicated to ensuring that general-purpose artificial intelligence benefits all of humanity. https://openai.com/ https://jobs.ashbyhq.com/openai/38ddaa2c-a490-427a-8457-0e92bf00138c San Francisco; New York City; Remote - US; Seattle 2026-04-24 0a12b071-6f4 Member of Technical Staff - Data Scientist As a data scientist at Microsoft AI, you will be tasked with helping us to maximize our GPU utilization while solving real issues in the AI frontier.

Our vision is bold and broad , to build systems that have true artificial intelligence across agents, applications, services, and infrastructure. It’s also inclusive: we aim to make AI accessible to all , consumers, businesses, developers , so that everyone can realize its benefits.

The abuse prevention team is responsible for identifying instances of abuse within copilot and blocking or mitigating those instances. You will be a critical component to keeping copilot safe and productive for all.

Microsoft’s mission is to empower every person and every organization on the planet to achieve more. As employees we come together with a growth mindset, innovate to empower others, and collaborate to realize our shared goals. Each day we build on our values of respect, integrity, and accountability to create a culture of inclusion where everyone can thrive at work and beyond.

Responsibilities: Drive product insights, opportunity analysis, and track metrics to support efforts across Microsoft Copilot. Drive new ways of instrumenting and measuring impact to evaluate new feature performance through experimentation. Define metrics and build basic data pipelines to enable A/B experimentation for new features and mitigating abusive users. Hands-on analysis of large volumes of telemetry data using various algorithms and tools including your own Articulate insights, storyboard with data and communicate to influence leadership and other key decision makers. Find a path to get things done despite roadblocks to get your work into the hands of users quickly and iteratively. Enjoy working in a fast-paced, design-driven, product development cycle. Work collaboratively with our engineers, Product Managers, and marketing to take ambiguous projects that drive user growth, engagement, and retention. This includes identifying market opportunities, optimizing app flows and improving product features and proposing innovation solutions based on data. Embody our Culture and Values.

XML job scraping automation by YubHub

]]> full-time staff hybrid $119,800 - $234,700 per year Data Science, Machine Learning, Artificial Intelligence, GPU Utilization, Telemetry Data Analysis, Experimentation, A/B Testing, Data Pipelines, Statistics, Mathematics, Cloud Computing, Big Data, Data Visualization, SQL, Python, R, Java, C++ Engineering Technology Microsoft AI https://logos.yubhub.co/microsoft.ai.png Microsoft AI is a subsidiary of Microsoft Corporation, a multinational technology company. https://microsoft.ai https://microsoft.ai/job/member-of-technical-staff-data-scientist-7/ Mountain View 2026-04-24