LeethubLeethub
JobsCompaniesBlog
Go to dashboard

Leethub

Curated tech jobs from FAANG and top companies worldwide.

Top Companies

  • Google Jobs
  • Meta Jobs
  • Amazon Jobs
  • Apple Jobs
  • Netflix Jobs
  • All Companies →

Job Categories

  • Software Engineering
  • Data, AI & Machine Learning
  • Product Management
  • Design & User Experience
  • Operations & Strategy
  • Remote Jobs
  • All Categories →

Browse by Type

  • Remote Jobs
  • Hybrid Jobs
  • Senior Positions
  • Entry Level
  • All Jobs →

Resources

  • Google Interview Guide
  • Salary Guide 2025
  • Salary Negotiation
  • LeetCode Study Plan
  • All Articles →

Company

  • Dashboard
  • Privacy Policy
  • Contact Us
© 2026 Leethub LLC. All rights reserved.
Home›Jobs›Baseten›Software Engineer - Model Performance
Baseten

About Baseten

Simplifying machine learning for every organization

🏢 Tech👥 21-100 employees📍 Union Square, San Francisco, CA💰 $285m
B2BAnalyticsBusiness IntelligenceMachine Learning

Key Highlights

  • Headquartered in Union Square, San Francisco, CA
  • $285 million raised in Series C funding
  • Team growth of 3x over the last five years
  • Unlimited PTO with a company-wide holiday break

Baseten is a machine learning application builder headquartered in Union Square, San Francisco, CA. With $285 million in funding from investors like Coatue Management and Founders Fund, Baseten simplifies AI integration for businesses, enabling data scientists to deploy ML models without needing spe...

🎁 Benefits

Baseten offers a remote-first work environment with a $1,000 stipend for home office setup, unlimited PTO with a company-wide break during the holiday...

🌟 Culture

Baseten's culture emphasizes simplifying complex AI technologies for businesses, fostering a collaborative environment where team members can connect ...

🌐 Website💼 LinkedIn𝕏 TwitterAll 39 jobs →
Baseten

Software Engineer - Model Performance

Baseten • San Francisco

Posted 1 year ago🏛️ On-SiteSoftware engineering📍 San francisco
Apply Now →

Job Description

ABOUT BASETEN

Baseten powers mission-critical inference for the world's most dynamic AI companies, like Cursor, Notion, OpenEvidence, Abridge, Clay, Gamma and Writer. By uniting applied AI research, flexible infrastructure, and seamless developer tooling, we enable companies operating at the frontier of AI to bring cutting-edge models into production. We're growing quickly and recently raised our $150M Series D, backed by investors including BOND, IVP, Spark Capital, Greylock, and Conviction. Join us and help build the platform engineers turn to to ship AI products.

THE ROLE

Are you passionate about advancing the application of artificial intelligence? We are looking for a Software Engineer focused on ML performance to join our dynamic team. This role is ideal for someone who thrives in a fast-paced startup environment and is eager to make significant contributions to the exciting field of LLM Inference. If you are a backend engineer who thrives on making things faster and is excited about open-source ML models, we look forward to your application.

EXAMPLE INITIATIVES

You'll get to work on these types of projects as part of our Model Performance team:

  • Baseten Embeddings Inference: The fastest embeddings solution available

  • The Baseten Inference Stack

  • Driving model performance optimization


RESPONSIBILITIES

  • Implement, refine, and productionize cutting-edge techniques (quantization, speculative decoding, kv cache reuse, chunked prefill and LoRA) for ML model inference and infrastructure.

  • Deep dive into underlying codebases of TensorRT, PyTorch, TensorRT-LLM, vllm, sglang, CUDA, and other libraries to debug ML performance issues.

  • Apply and scale optimization techniques across a wide range of ML models, particularly large language models.

  • Collaborate with a diverse team to design and implement innovative solutions.

  • Own projects from idea to production.

REQUIREMENTS

  • Bachelor's, Master's, or Ph.D. degree in Computer Science, Engineering, Mathematics, or related field.

  • Experience with one or more general-purpose programming languages, such as Python or C++.

  • Familiarity with LLM optimization techniques (e.g., quantization, speculative decoding, continuous batching).

  • Strong familiarity with ML libraries, especially PyTorch, TensorRT, or TensorRT-LLM.

  • Demonstrated interest and experience in LLM’s.

  • Deep understanding of GPU architecture.

  • Bonus:

    • Proficiency in enhancing the performance of software systems, particularly in the context of large language models (LLMs).

    • Experience with CUDA or similar technologies.

    • Deep understanding of software engineering principles and a proven track record of developing and deploying AI/ML inference solutions.

    • Experience with Docker and Kubernetes.

BENEFITS

  • Competitive compensation, including meaningful equity.

  • 100% coverage of medical, dental, and vision insurance for employee and dependents

  • Generous PTO policy including company wide Winter Break (our offices are closed from Christmas Eve to New Year's Day!)

  • Paid parental leave

  • Company-facilitated 401(k)

  • Exposure to a variety of ML startups, offering unparalleled learning and networking opportunities.

Apply now to embark on a rewarding journey in shaping the future of AI! If you are a motivated individual with a passion for machine learning and a desire to be part of a collaborative and forward-thinking team, we would love to hear from you.

At Baseten, we are committed to fostering a diverse and inclusive workplace. We provide equal employment opportunities to all employees and applicants without regard to race, color, religion, gender, sexual orientation, gender identity or expression, national origin, age, genetic information, disability, or veteran status.

Interested in this role?

Apply now or save it for later. Get alerts for similar jobs at Baseten.

Apply Now →Get Job Alerts