추천 검색

Software Engineer - GPU Kernel
공고 원문
주요업무
About the job FriendliAI is looking for a GPU Kernel Engineer to design, build, and optimize the low-level compute kernels that power our large-scale, GPU-accelerated AI inference platform. You will be delivering world-class inference speed across NVIDIA and AMD GPUs. With our recent $20M funding, we are scaling our team to meet market demand. This is a deeply technical, high-impact role where you will write GPU code, implement advanced optimizations. As part of our engine team, you will contribute directly to the company's proprietary inference engine which supports over 450,000 models on Hugging Face. You will work with the inventors of continuous batching and collaborate with the platform team to deploy your work into production. Key Responsibilities Design, implement, and optimize high-performance GPU kernels for AI inference (e.g., GEMM, attention, routing) Develop and maintain GPU code in CUDA and C++, including low-level assembly when needed Implement reduced-precision and quantized kernels (FP8/FP4) for low-latency or high-throughput inference Benchmark and ensure cross-vendor performance parity between NVIDIA and AMD hardware Contribute to internal GPU libraries and tune performance of performance-critical components Accelerate multi-modal model pipelines Investigate and integrate next-generation GPU features
자격요건
3+ years of experience in GPU programming, HPC, or performance-critical systems Bachelor's or Master's degrees in Computer Science, Computer Engineering, Electrical Engineering, or a related field Strong proficiency in CUDA for NVIDIA GPUs or ROCm/HIP for AMD GPUs Deep understanding of GPU architecture: warps, threads, memory hierarchy, synchronization, and latency-throughput trade-offs Proficiency in C++ Experience with GPU profiling and performance tuning Strong numerical background with understanding of precision trade-offs and quantization techniques
우대사항
Experience optimizing transformer, multi-modal, or Mixture-of-Experts (MoE) architectures at the kernel level Familiarity with the latest GPU libraries and frameworks (CUTLASS, Triton, …) Inter-GPU communication programming experience Open-source contributions related to GPU performance or ML acceleration Research or conference presentations on GPU optimization, HPC, or numerical computing
기술스택
C++, Python, CUDA
채용절차
서류 전형 → 1 차 인터뷰 → 2 차 인터뷰 → 컬처핏 인터뷰 → 최종 합격 근무 형태: 정규직 근무 환경: 유연한 원격 근무 및 필요시 오피스 출근을 병행하는 하이브리드 근무 Aggressive Compensation: 업계 최고 수준의 기본급과 성과에 따른 상한선 없는(Uncapped) 파격적인 커미션 구조 제공 Stock Option: 회사의 성장과 함께 자산 증식이 가능한 스톡옵션 부여 채용 절차는 상황에 따라 일부 변경될 수 있습니다
이런 공고는 어때요?

메텍홀딩스
10년 이하 · 정규직 · 서울

클루닉스
5년~20년 · 정규직 · 서울

써냅스
10년 이하 · 정규직 · 서울

님버스테크
20년 이하 · 정규직 · 대구 외 2개

제너레잇
3년~20년 · 정규직 · 서울

타인에이아이
10년 이하 · 정규직 · 서울

코리아포트원
5년~20년 · 정규직 · 서울

피엔에스테크놀러지
20년 이하 · 정규직 · 경기

타이드풀
5년~10년 · 정규직 · 경기

디사일로
20년 이하 · 체험형인턴 · 서울

팀카이
20년 이하 · 체험형인턴 · 서울

알엑스
5년~10년 · 정규직 · 서울

위시켓
3년~10년 · 정규직 · 서울
.webp)
토모도모
3년~20년 · 정규직 · 서울

AMD
3년 이하 · 정규직 · 해외

현대자동차
5년 이상 · 정규직 · 서울

AWS
10년 이상 · 정규직 · 서울

CJ대한통운
3년 이상 · 정규직 · 서울

아마존
2년 이상 · 정규직 · 서울

AMD
3년 이하 · 체험형인턴 · 해외




