LeethubLeethub
JobsCompaniesBlog
Go to dashboard

Leethub

Curated tech jobs from FAANG and top companies worldwide.

Top Companies

  • Google Jobs
  • Meta Jobs
  • Amazon Jobs
  • Apple Jobs
  • Netflix Jobs
  • All Companies →

Job Categories

  • Software Engineering
  • Data, AI & Machine Learning
  • Product Management
  • Design & User Experience
  • Operations & Strategy
  • Remote Jobs
  • All Categories →

Browse by Type

  • Remote Jobs
  • Hybrid Jobs
  • Senior Positions
  • Entry Level
  • All Jobs →

Resources

  • Google Interview Guide
  • Salary Guide 2025
  • Salary Negotiation
  • LeetCode Study Plan
  • All Articles →

Company

  • Dashboard
  • Privacy Policy
  • Contact Us
© 2026 Leethub LLC. All rights reserved.
Home›Jobs›Nebius AI›Technical Product Manager (Cluster Experience)
Nebius AI

About Nebius AI

Empowering AI with robust infrastructure solutions

🏢 Tech👥 51-250📅 Founded 2022📍 Amsterdam, North Holland, Netherlands

Key Highlights

  • Publicly traded on Nasdaq, expanding AI infrastructure market
  • Headquartered in Amsterdam with hubs in the US, Europe, and Israel
  • Team of around 400 skilled engineers focused on AI/ML
  • Specializes in large-scale GPU clusters and cloud platforms

Nebius is a Nasdaq-listed company headquartered in Amsterdam, specializing in AI infrastructure solutions. With a team of around 400 engineers, Nebius provides large-scale GPU clusters and cloud platforms designed to support the rapid growth of the AI industry. The company has established R&D and co...

🎁 Benefits

Nebius offers competitive equity packages, a flexible PTO policy, and opportunities for remote work. Employees also benefit from a learning budget to ...

🌟 Culture

Nebius fosters a culture centered around engineering excellence and innovation in AI infrastructure. The company values collaboration across its globa...

🌐 Website💼 LinkedInAll 179 jobs →
Nebius AI

Technical Product Manager (Cluster Experience)

Nebius AI • Amsterdam, Netherlands; Berlin, Germany; France; Netherlands; Prague, Czech Republic; Remote - Europe

Posted 1w ago🏠 RemoteMid-LevelTechnical product manager📍 Amsterdam📍 Berlin📍 Prague
Apply Now →

Skills & Technologies

Machine learningDistributed systemsCloud engineering

Job Description

Why work at Nebius
Nebius is leading a new era in cloud computing to serve the global AI economy. We create the tools and resources our customers need to solve real-world challenges and transform industries, without massive infrastructure costs or the need to build large in-house AI/ML teams. Our employees work at the cutting edge of AI cloud infrastructure alongside some of the most experienced and innovative leaders and engineers in the field.

Where we work
Headquartered in Amsterdam and listed on Nasdaq, Nebius has a global footprint with R&D hubs across Europe, North America, and Israel. The team of over 800 employees includes more than 400 highly skilled engineers with deep expertise across hardware and software engineering, as well as an in-house AI R&D team.

The role

We're building the leading platform for large-scale Machine Learning (ML) training and inference — powering workloads from a few nodes to thousands of GPUs. We're looking for a Product Manager who will define how customers experience GPU clusters: their reliability, performance, and overall usability at scale. As a Product Manager in the Cluster Experience team, you will own foundational tracks that shape how ML teams train and serve models on multi-node distributed systems. Your initial focus will be reliability, performance, and observability for large-scale training and distributed inference. Over time, the role expands into UX, operational tooling, and advanced cluster workflows. This is a deeply technical PM role, but it does not require prior product management experience - strong candidates with backgrounds in ML infrastructure, distributed systems, SRE, or cloud engineering who want to grow into product are welcome. If you want to influence how state-of-the-art models are trained and deployed at scale, this role is for you.

Your responsibilities will include: 

  • Own key tracks in Cluster Experience: reliability, performance, and user experience for distributed ML workloads.
  • Define product direction from problem discovery → design → delivery → adoption, working closely with engineering and research teams.
  • Drive cross-functional execution across compute, networking, storage, observability, and platform teams.
  • Perform deep customer research: interviews, analytics, and workload studies to identify bottlenecks across hardware, network, scheduler, and runtime.
  • Translate state-of-the-art ML papers ideas into practical, scalable product features for large GPU clusters.
  • Shape how users interact with clusters - from dashboards and notifications to partitioning, node management, and training observability.  

We expect you to have: 

  • 3–5+ years of experience in product management, ML infrastructure/MLOps, distributed systems engineering, or cloud architecture.
  • Strong technical foundation in computer science, distributed systems, or ML infrastructure.
  • Hands-on familiarity with ML training, ideally using orchestrators like Slurm, Kubernetes, Ray, or similar systems.
  • Proven ability to ship technically complex features with multiple engineering teams.
  • Excellent communicator capable of influencing engineering, research, and customer stakeholders.
  • Experience with product analytics, data-driven prioritization, and experiment design.
  • Strong willingness and ability to learn quickly in a fast-evolving ML and infrastructure environment.

It will be an added bonus if you have: 

  • Experience working with GPU platforms, Infiniband/RDMA networking, or HPC systems.
  • Understanding of modern ML frameworks (PyTorch, DeepSpeed, FSDP, NCCL, etc.).
  • Knowledge of ML training efficiency: Goodput, MFU, scheduling, health checks.
  • Exposure to LLM training, distributed data/ZeRO/FSDP strategies, or transformer inference.
  • Experience in observability, performance tuning, or reliability engineering.
  • Customer-facing technical experience (supporting ML or infrastructure workloads).

About Nebius

Nebius AI is an AI cloud platform with one of the largest GPU capacities in Europe. Launched in November 2023, the Nebius AI platform provides high-end, training-optimized infrastructure for AI practitioners. As an NVIDIA preferred cloud service provider, Nebius AI offers a variety of NVIDIA GPUs for training and inference, as well as a set of tools for efficient multi-node training. 

Nebius AI owns a data center in Finland, built from the ground up by the company’s R&D team and showcasing our commitment to sustainability. The data center is home to ISEG, the most powerful commercially available supercomputer in Europe and the 16th most powerful globally (Top 500 list, November 2023).  

Nebius’s headquarters are in Amsterdam, Netherlands, with teams working out of R&D hubs across Europe and the Middle East. 

Nebius AI is built with the talent of more than 500 highly skilled engineers with a proven track record in developing sophisticated cloud and ML solutions and designing cutting-edge hardware. This allows all the layers of the Nebius AI cloud – from hardware to UI – to be built in-house, distictly differentiating Nebius AI from the majority of specialized clouds: Nebius customers get a true hyperscaler-cloud experience tailored for AI practitioners. We’re growing and expanding our products every day. 

If you’re up to the challenge and are excited about AI and ML as much as we are, join us!

What we offer 

  • Competitive salary and comprehensive benefits package.
  • Opportunities for professional growth within Nebius.
  • Flexible working arrangements.
  • A dynamic and collaborative work environment that values initiative and innovation.

We’re growing and expanding our products every day. If you’re up to the challenge and are excited about AI and ML as much as we are, join us!

Interested in this role?

Apply now or save it for later. Get alerts for similar jobs at Nebius AI.

Apply Now →Get Job Alerts