GPU Performance Engineer

genmoยท Engineering
Apply Now โ†—
๐Ÿ“ San Francisco HQFullTime

About this role

We are Genmo, a research lab dedicated to building open, state-of-the-art models for video generation towards unlocking the right brain of AGI. Join us in shaping the future of AI and pushing the boundaries of what's possible in video generation.

We're seeking a GPU Performance Engineer to squeeze every last FLOP from our H100 infrastructure and optimize our model serving stack to its absolute limits.

The Role

You'll be our performance optimization expert, using advanced profiling tools to identify bottlenecks and implementing solutions that achieve 5-10x speedups. From writing custom CUDA kernels to eliminating cold start latency, you'll ensure our infrastructure delivers world-class performance. This role is perfect for someone who gets excited about microsecond optimizations and pushing hardware to its theoretical limits.

Key Responsibilities

  • Profile and optimize GPU workloads using Nsight Systems, nvprof, and custom instrumentation

  • Write high-performance CUDA and Triton kernels for critical model operations

  • Optimize cold start latency from seconds to milliseconds for our serving infrastructure

  • Tune memory access patterns, kernel fusion, and GPU utilization

  • Collaborate with ML engineers to optimize model implementations

  • Debug performance issues across the full stack from application to hardware

  • Implement custom memory pooling and allocation strategies

  • Share optimization techniques and build performance culture across teams

Qualifications

  • Bachelor's or Master's degree in Computer Science, Electrical Engineering, or related field

  • 5+ years systems programming experience with 3+ years focused on GPU optimization

  • Expert proficiency with GPU profiling tools (Nsight Systems, nvprof)

  • Strong CUDA programming skills with production kernel development

  • Deep understanding of GPU architecture (memory hierarchy, SMs, warps)

  • Track record of achieving significant performance improvements (5-10x)

  • Experience with Python and C++ in production environments

We Value

  • Experience with Triton kernel development

  • Knowledge of CUTLASS or similar high-performance libraries

  • Background in ML-specific optimizations (attention, transformers)

  • RDMA/InfiniBand optimization experience

  • Contributions to GPU libraries or frameworks

  • Low-level debugging skills (PTX/SASS reading)

Genmo is an Equal Opportunity Employer. Candidates are evaluated without regard to age, race, color, religion, sex, disability, national origin, sexual orientation, veteran status, or any other characteristic protected by federal or state law. Genmo, Inc. is an E-Verify company and you may review the Notice of E-Verify Participation and the Right to Work posters in English and Spanish.

Frequently Asked Questions

Is the salary disclosed for the GPU Performance Engineer position at genmo?
The salary for this GPU Performance Engineer role at genmo is not publicly listed. Click "Apply Now" to learn more about the compensation package on their official careers page.
Where is the GPU Performance Engineer position at genmo located?
This GPU Performance Engineer role at genmo is based in San Francisco HQ. The position is listed as on-site or hybrid. Check the full job description or apply directly to confirm the work arrangement.
Is the GPU Performance Engineer role at genmo full-time or part-time?
This is listed as a FullTime position. It is posted as a GPU Performance Engineer role in the Engineering department at genmo.
Which team or department does the GPU Performance Engineer at genmo belong to?
This GPU Performance Engineer position is part of the Engineering department at genmo. See the full job description for more information about the team structure and responsibilities.
How do I apply for the GPU Performance Engineer position at genmo?
Click the "Apply Now" button on this page. You will be redirected to genmo's official application portal hosted on ashby where you can submit your application directly.
When was the GPU Performance Engineer job at genmo posted?
This GPU Performance Engineer position at genmo was posted on Jul 17, 2025. Apply as soon as possible โ€” early applications are often reviewed first.
GPU Performance Engineer
genmo
Apply for this role โ†—

You'll be redirected to genmo's official application page on Ashby ATS.