추천 검색

Solution Architect - AI Inference Specialist
공고 원문
주요업무
About the job FriendliAI is seeking a Forward Deployed Engineer (FDE) to assist enterprises in deploying, scaling, and operating generative and agentic AI workloads on FriendliAI infrastructure. You will work directly with customers to solve and implement production-grade applications using our products, such as Serverless Endpoints, Dedicated Endpoints, or Container. Friendli Container is our service that allows customers to download our inference engine as Docker images and deploy it in their chosen environment, such as private clouds or on-premises. Our Friendli Container can be adopted directly to AWS EKS clusters using our EKS add-on product. You will work directly on our customers’ projects, collaborating with their engineering teams to solve AI inference challenges like scaling, orchestration, and monitoring. This is a hands-on, customer-embedded role. If you have worked in DevOps, platform engineering, or SRE for AI applications, this is your ideal position. Key Responsibilities Design and implement large-scale deployment architectures for LLM and multimodal inference Deploy and manage containerized workloads across Kubernetes clusters Diagnose production issues, such as performance bottlenecks, and implement temporary fixes as needed Collaborate with customers’ DevOps teams to integrate FriendliAI’s infrastructure into their CI/CD workflows Develop scripts, Helm charts, and Terraform modules that simplify repeated deployments Contribute field insights to shape our platform reliability, observability, and scaling strategies Lead workshops, technical sessions, or webinars to help customers master infrastructure best practices
자격요건
Qualifications 3+ years of experience in cloud infrastructure, DevOps, or reliability engineering Bachelor’s or Master's degree in Computer Science, Computer Engineering, Electrical Engineering, or equivalent Proficiency with Kubernetes, Docker, Terraform, and Helm Strong foundation in distributed systems, networking, and performance tuning Experience with GPU-based computing and generative AI model serving workloads Strong technical background in backend systems or AI tooling Experience operating workloads on AWS, GCP, or OCI Excellent problem-solving and debugging skills in real-world environments
우대사항
Preferred Experience Experience deploying large models (LLMs, diffusion models) on GPUs or clusters Familiarity with inference frameworks (Triton, vLLM, TensorRT, DeepSpeed-Inference) Familiarity with observability stacks (Prometheus, Grafana, Loki, ELK, OTEL) Understanding of networking security and compliance frameworks (e.g., SOC 2) Experience supporting on-prem or hybrid-cloud deployments
기술스택
Docker, Helm, Kubernetes, Terraform, AWS-EKS
채용절차
서류 전형 → 1 차 인터뷰 → 2 차 인터뷰 → 컬처핏 인터뷰 → 최종 합격 근무 형태: 정규직 근무 환경: 유연한 원격 근무 및 필요시 오피스 출근을 병행하는 하이브리드 근무 Aggressive Compensation: 업계 최고 수준의 기본급과 성과에 따른 상한선 없는(Uncapped) 파격적인 커미션 구조 제공 Stock Option: 회사의 성장과 함께 자산 증식이 가능한 스톡옵션 부여 채용 절차는 상황에 따라 일부 변경될 수 있습니다
이런 공고는 어때요?

신화씨엔에스
20년 이하 · 정규직 · 강원
.jpg)
에이블네트워크
10년 이하 · 정규직 · 경기

엘리스그룹
5년~10년 · 정규직 · 서울

미리비트
5년~10년 · 정규직 · 경기

vox.ai
1년~9년 · 정규직 · 서울

몬드리안에이아이
5년~20년 · 정규직 · 인천

피트인
7년~10년 · 정규직 · 기타

와탭랩스
6년~10년 · 정규직 · 서울

더페어글로벌
4년~10년 · 정규직 · 서울

미리디
3년~6년 · 정규직 · 서울

블리츠다이나믹스
3년~10년 · 정규직 · 서울
.png)
윤회
3년~10년 · 정규직 · 서울

신화씨엔에스
10년 이하 · 정규직 · 서울

님버스테크
20년 이하 · 정규직 · 대구 외 2개

엘리스그룹
10년 이하 · 정규직 · 서울

메크랩
3년~10년 · 정규직 · 기타

엘리스그룹
3년~20년 · 정규직 · 서울

스콘에이아이(SconAI)
1년~3년 · 정규직 · 서울
프라임엔시스템
1년~20년 · 정규직 · 서울

페이타랩
10년 이하 · 정규직 · 서울



