転職AIログインマイページ求人へ
MLOpsエンジニア条件を変える閉じる

こだわり
条件をクリア

ML Operations Engineer (AI/LLM)

株式会社メルカリ

テック 記載なし 企業サイト 掲載 9/2

給与

記載なし※ 募集要項に金額の記載がありません

勤務地
記載なし
雇用形態
正社員
働き方
記載なし
年間休日
記載なし

要約求人票をもとにAIがまとめたものです

As an MLOps engineer on the AI/LLM team, you will own production serving, deployment, and operations for machine learning and LLM models in a cloud-native environment. The role includes model inference orchestration, serving and deployment, performance optimization, monitoring and reliability, model quality evaluation, and enabling research-to-production workflows. The platform serves tens of millions of users.

応募資格

必須

  • Bachelor's degree in Computer Science, Engineering, or a related field, or equivalent practical experience.
  • 5+ years of software engineering experience, including proven experience in production MLOps: end-to-end model deployment, serving, and CI/CD in cloud environments.
  • Experience designing and operating large-scale, high-availability distributed systems, including observability, SLO definition, and incident response.
  • Strong experience in cloud-native infrastructure (Kubernetes, Docker).
  • Proficiency in Python and infrastructure-as-code (Terraform).
  • Excellent written and verbal communication.
  • English: Proficient (CEFR - B2)

歓迎

  • Experience integrating ML serving with large-scale distributed data layers (e.g., data warehouses, wide-column stores, in-memory caches).
  • Expertise in model inference optimization (TensorRT-LLM, quantization, JAX).
  • Experience operating large-scale model inference gateways and orchestrators.
  • 2+ years of hands-on experience operating GenAI/LLM workloads in production (e.g., LLM serving frameworks, token throughput and cost optimization).
  • Experience building LLM evaluation, guardrail, or quality-monitoring pipelines (e.g., LLM-as-judge, golden datasets, drift detection).
  • Experience with serving infrastructure for RAG or agentic AI workloads (vector search, tool-calling execution environments).
  • Experience partnering closely with research or data science teams to bring research innovations into production.
  • Master's or Ph.D. in a related technical field.
  • Japanese: Independent (CEFR - B2) optional

使用ツール

DataServBigQueryBigTableValkeyModel Inference GatewayConsoleNVIDIATPUTriton Inference ServerJAX/TPU Gatewaysmodel repositoriesdynamic batchingconcurrent model executionCI/CDTerraformKubernetesKV caching

募集要項最終確認 9/26

職種
MLOpsエンジニア
雇用形態
Employment Status: Full-time
給与
記載なし
勤務地
Roppongi
働き方
記載なし
勤務時間
Work Hours: Full Flextime (no core time)
経験年数
5年以上
学歴
大学卒業以上
選考の流れ
Application screeningSkill assessment: For engineering positions, you will be asked to complete a skill assessment on HackerRank or GitHub. For non-engineering positions, you may be asked to complete an assessment depending on the position. (The timing of the assessment may coincide with the interview process.)Interview: The number of interviews may vary depending on the position.Reference check: We will ask for online references around the timing of the final interview.Offer: Offers will be determined carefully in consideration of the final interview and the reference check.
掲載日
2026/09/02 18:20(27日前)
最終確認
2026/09/26 01:59(3日前)

仕事内容

Organization/Team Mission

The AI / LLM Team’s mission is focused on three core pillars, “product”, "enablement" and "research", delivering new AI-driven features and user experiences to maximize product-facing impact for Mercari's business.

We do this both through independent initiatives owned by our team, as well as by horizontally collaborating with product, engineering, and research teams across the entire organization.

As an MLOps engineer on the AI/LLM team, you will own how our machine learning and LLM models reach production and stay healthy there in our cloud-native environment. Your focus is the production serving, deployment, and operations that turn models into reliable, cost-efficient services, seamlessly integrating with our machine learning operations to serve tens of millions of users.

Work Responsibilities

Data and Model Orchestration: Own the end-to-end orchestration of model inference, including integrating with DataServ for retrieval (BigQuery, BigTable, Valkey) and managing the Model Inference Gateway and Console.

Model Serving and Deployment: Own production model serving on the cloud-native NVIDIA and TPU stacks (Triton Inference Server, TensorRT-LLM, JAX/TPU Gateways). Manage model repositories, dynamic batching, and concurrent model execution. Build CI/CD, rollout, and rollback paths for safe, scalable model deployment, and automate provisioning and lifecycle management (Terraform, Kubernetes) so the platform scales seamlessly across teams.

Inference Performance Optimization: Profile and optimize deployments across LLM and non-LLM workloads, including model compilation, quantization, and batching strategies to hit latency and throughput targets while managing cost. Maintain performance baselines and regression detection.

Monitoring and Reliability: Build robust monitoring and alerting for model health and latency, including service-level metrics for the data retrieval and inference gateway layers. Define SLOs and own on-call and incident response for the serving layer.

Model Quality and Evaluation: Build automated evaluation and quality monitoring into the deployment path, regression and drift detection, offline/online evaluation, and LLM output quality checks, so models stay healthy long after launch.

Research-to-Production Enablement: Partner with ML engineers and researchers to turn experimental models into production-ready services, providing self-service workflows and abstractions that let teams deploy safely and quickly without deep infrastructure expertise.

Unique Challenges

Build and operate the production ML serving and Data Orchestration platform behind Mercari Group's AI and LLM features, serving tens of millions of users.

Drive the strategy for model inference at scale, bridging the gap between complex data retrieval and fast-moving ML model inference to ensure high-performance, cost-effective service delivery.

Shape Mercari's next-generation LLM serving stack, from inference optimization (quantization, dynamic batching, KV caching) to the evaluation and execution infrastructure needed for emerging agentic AI workloads.

この求人は株式会社メルカリの採用ページの掲載内容をもとに構成しています。応募条件の最新情報は募集元をご確認ください。

募集元の採用ページから応募できます

応募は株式会社メルカリの採用ページで受け付けています。このページは採用ページの掲載内容をもとに構成しているため、最新の応募条件は募集元でご確認ください。

募集元の採用ページを開く
相談募集元で応募する採用ページへ

AI相談