인기 검색

Senior Site Reliability Engineer· Cancer Screening
공고 요약
담당업무
Lunit INSIGHT 제품 및 관련 서비스의 안정성·가용성·성능·운영 품질 개선
Azure 클라우드 인프라, 배포, 모니터링, 운영 자동화 시스템 설계 및 개선
클라우드 리소스 사용량·비용 분석과 안정성·성능을 고려한 비용 최적화
모니터링·알림·런북·자동화를 통한 장애 진단, 복구, 근본 원인 분석 및 재발 방지
환경·고객별 설정, 인증서, 시크릿, 접근 권한 등 운영 구성과 보안 요소 관리
제품 및 개발팀과 협업해 설계·개발 단계부터 안정성 반영
Lunit International SRE 팀과 영어로 협업하며 글로벌 운영 환경과 국내 제품 개발 환경 정렬
24시간 서비스 운영을 위한 온콜·인시던트 대응 체계 구축 및 온콜 로테이션 참여
자격요건
SRE, DevOps, 플랫폼 엔지니어링 또는 프로덕션 인프라 운영 경험
Azure 기반 프로덕션 환경을 직접 설계·운영한 경험
Linux, 네트워킹, 컨테이너 실무 경험
CI/CD, Infrastructure as Code 또는 운영 자동화 시스템 설계·개선 경험
모니터링 기반 프로덕션 장애 분석, 복구 및 재발 방지 활동 경험
제품·개발팀과 협업해 안정성과 개발 속도의 균형을 조정한 경험
해외 엔지니어 및 다양한 직군과 기술적 의사결정을 영어로 논의·조율하는 능력
Python, Bash, Azure, Linux, Docker, Kubernetes, Terraform, Bicep 사용 환경
GitHub Actions, Azure DevOps Pipelines, Azure Monitor, Application Insights, Log Analytics, PagerDuty, ServiceNow 활용 환경
Python, Go, FastAPI, PostgreSQL, Redis, RabbitMQ 및 Jira, Confluence, Slack, GitHub 활용 환경
우대사항
외국어 능통자 우대
영어 기반 글로벌 SRE 협업
복리후생
건강검진 지원
자기계발 지원
마감기한
상시채용
루닛 지원하기 전 이력서 Check
시간이 없다면 AI와 한번에 수정해보세요
공고 원문
소개
Lunit, a portmanteau of ‘Learning unit,’ is a medical AI software company devoted to providing AI-powered total cancer care. Our AI solutions help discover cancer and predict cancer treatment outcomes, achieving timely and individually-tailored cancer treatment. [About the Team] • Lunit's Software Engineering (hereafter, "SE") department develops the Lunit INSIGHT product. Within the SE department, the Core Product Engineering team performs the Backend/Application development needed to productize INSIGHT AI models. We translate product requirements into development requirements and build Software as a Medical Device (SaMD) that complies with medical device guidelines. • Our team's core goal is to ensure the INSIGHT product has optimized performance and reliability so it can operate across diverse countries and deployment environments. [About the Position] • The Site Reliability Engineer role sits between development and operations, improving reliability, observability, automation, and operational systems so the INSIGHT product can run stably. • Rather than simply performing operational tasks, you will discover recurring problems and operational inefficiencies in production, analyze their root causes, and translate the findings into automation, improved observability, and improved operational processes. • You will collaborate with Lunit International's SRE team, understand differing operational environments and processes, and connect and align the global operating system with the domestic product development environment.
주요업무
You will design and improve the reliability, observability, deployment automation, and cloud infrastructure operations of the Lunit INSIGHT product to ensure it runs stably. This role is not limited to performing predefined operational tasks. You will identify gaps in the team's current SRE capabilities and operational systems, then define and drive the improvements needed. • Improve the stability, availability, performance, and operational quality of the Lunit INSIGHT product and related services. • Design and enhance cloud infrastructure, deployment, monitoring, and operational automation systems. •Analyze cloud resource usage and cost, and continuously drive cost optimization while considering stability and performance. • Diagnose, recover from, and perform root-cause analysis on incidents, and prevent recurrence through monitoring, alerting, runbooks, and automation. • Reliably manage and improve operational configurations and security elements such as per-environment/per-customer settings, certificates, secrets, and access permissions. • Understand the product domain and collaborate with development teams to build reliability in from the design and development stages. • Collaborate in English with Lunit International's SRE team, understanding and aligning differing operational environments and processes. • Build an on-call and incident-response system for 24/7 service operations, and participate in the actual on-call rotation to handle production incidents. As the team grows, evolve and lead a sustainable on-call operating model. [Tech Stack] • Scripting/Language: Python, Bash • Cloud/Infrastructure: Azure, Linux, Docker, Kubernetes, Terraform, Bicep • CI/CD: GitHub Actions, Azure DevOps Pipelines • Observability/Operations: Azure Monitor, Application Insights, Log Analytics, PagerDuty, ServiceNow, Custom Dashboards • Product Environment: Python, Go, FastAPI, PostgreSQL, Redis, RabbitMQ • Collaboration: Jira, Confluence, Slack, GitHub
자격요건
• 5+ years of experience in SRE, DevOps, Platform Engineering, or production infrastructure operations • Experience directly designing and operating Azure-based production environments • Hands-on experience with Linux, networking, and containers • Experience designing or improving CI/CD, Infrastructure as Code, or operational automation systems • Experience analyzing production incidents based on monitoring, and performing recovery and recurrence-prevention activities • Experience collaborating with product and development teams to balance stability and development velocity • Ability to discuss and coordinate technical decisions fluently in English with overseas engineers and colleagues from diverse roles
우대사항
• Experience establishing and improving the technical direction or operating systems of the SRE or Platform domain • Experience identifying the causes of recurring incidents or operational inefficiencies and leading structural improvements such as recurrence prevention or automation • Experience building or operating SRE practices such as on-call, incident management, postmortems, and SLO/SLA • Experience with Infrastructure as Code (Terraform, Bicep, ARM templates) and operational automation using Python, Bash, etc. • Experience operating container and deployment environments such as Kubernetes, GitHub Actions, and Azure DevOps • Experience using observability tools such as Azure Monitor, Application Insights, and Log Analytics, and incident-response/operations tools such as PagerDuty and ServiceNow
혜택 및 복지
• The office is at a very convenient location, just a minute away from Gangnam Station Exit 3. • Meal Allowance is provided (up to 12,000 KRW per meal) when working at the office. • Latest computer models, such as Macs and 4K monitors are provided and can be renewed every three years. • Seminar registration fees and book purchases are covered. • Regular in-house AI and medical seminars are held. • In-house English lessons (aka Luniversal) are provided for English development. • Access to high-quality AI learning resources & deep learning DevOps system. • Up to 1.2 million KRW worth of benefits points can be claimed annually. • Holiday Allowances are provided in the form of gifts or vouchers for Korean National holidays, Seollal and Chuseok. • Congratulatory and Condolence allowances, along with paid time off are provided. • Annual medical checkups and employee accident insurance are provided.
댓글
루닛에 맞는
이력서로 자동 수정하기