추천 검색

Software Engineer - AI Inference System
공고 원문
주요업무
About the job We are seeking a highly technical Inference Engine Engineer to optimize the performance and efficiency of our core inference engine. You will focus on designing, implementing, and optimizing GPU kernels and supporting infrastructure for next-generation generative and agentic AI workloads. Your work will directly power the most latency-critical and compute-intensive systems deployed by our customers. The ideal candidate is an exceptional engineer with a strong foundation in GPU programming and compiler infrastructure. You enjoy pushing the performance boundaries and have experience supporting production-scale machine learning applications. Key Responsibilities Design and optimize custom GPU kernels for AI (e.g., transformer and diffusion) workloads Contribute to the development of FriendliAI's kernel compiler, memory planner, runtime, and other core components Collaborate with cloud and infrastructure engineers to ensure end-to-end inference performance Analyze performance bottlenecks across the software and hardware stack, and implement targeted optimizations Drive support for new model architectures and tensor compute patterns Maintain production-grade performance infrastructure, including profiling, benchmarking, and validation tools
자격요건
Qualifications 5+ years of experience in production or high-impact research environments Production-level expertise in Python and C++ Bachelor's or Master's degree in Computer Science, Computer Engineering, Electrical Engineering, or equivalent Experience developing machine learning frameworks or performance-critical runtime systems Hands-on experience writing and optimizing GPU kernels Hands-on experience profiling GPU kernels Experience working with generative AI models such as transformer and diffusion models
우대사항
Preferred Experience Experience developing machine learning compilers or code generation systems Familiarity with dynamic shape compilation, memory planning, and kernel fusion Contributions to inference engines, compilers, or high-performance numerical libraries Understanding of multi-GPU and distributed inference strategies
기술스택
C++, Python
채용절차
서류 전형 → 1 차 인터뷰 → 2 차 인터뷰 → 컬처핏 인터뷰 → 최종 합격 근무 형태: 정규직 근무 환경: 유연한 원격 근무 및 필요시 오피스 출근을 병행하는 하이브리드 근무 Aggressive Compensation: 업계 최고 수준의 기본급과 성과에 따른 상한선 없는(Uncapped) 파격적인 커미션 구조 제공 Stock Option: 회사의 성장과 함께 자산 증식이 가능한 스톡옵션 부여 채용 절차는 상황에 따라 일부 변경될 수 있습니다
이런 공고는 어때요?

트레독스
3년~10년 · 정규직 · 서울

스몰티켓
7년~15년 · 정규직 · 서울

큐픽스
3년~8년 · 정규직 외 1개 · 경기

두노소프트
9년~20년 · 계약직 외 1개 · 서울

YesPlz AI
3년~8년 · 정규직 · 기타

블랙피그에이아이
3년~7년 · 정규직 · 서울

르몽
7년~20년 · 정규직 · 서울

엘리스그룹
5년~10년 · 정규직 · 서울

이노바이드
5년~20년 · 정규직 · 서울

몬드리안에이아이
2년~20년 · 정규직 · 인천

키즐링
4년~10년 · 정규직 · 서울

인포뱅크
4년~8년 · 정규직 · 경기

재미스튜디오
10년 이하 · 정규직 · 대전

펄크럼테크놀로지스
3년~20년 · 정규직 · 서울

두노소프트
3년~20년 · 계약직 외 1개 · 서울

벙커키즈
20년 이하 · 정규직 · 서울

지엔에이컴퍼니
2년~10년 · 정규직 · 서울

어쎈드(ASCEND)
10년 이하 · 정규직 · 서울

모다플
3년~10년 · 정규직 · 서울

래빗홀컴퍼니
10년 이하 · 정규직 외 1개 · 서울




