Search by job, company or skills

Head of LLM Evaluation

10-12 Years
  • Posted a day ago
  • Be among the first 10 applicants

Job Description

Key Responsibilities

  1. Define Model Evaluation Strategy
  • Design a Comprehensive Evaluation Framework: Develop and maintain a standardized evaluation framework that measures model performance across every stage of the LLM lifecycle, from pre-training to production deployment.
  • Define Quality Metrics and Success Criteria: Establish clear Key Performance Indicators (KPIs) and acceptance thresholds that determine whether a model is ready to progress to the next development stage or production deployment.
  • Develop Indonesia-Centric Evaluation Benchmarks: Create and maintain benchmark datasets that accurately measure the model's understanding of Bahasa Indonesia, regional dialects, cultural context, local regulations, and industry-specific terminology.
  • Define Quality Gates: Define the evaluation process that every model must successfully complete before being approved for deployment.

Benchmark Development

  • Design Comprehensive Benchmark Suites: Develop and maintain standardized benchmark datasets that evaluate the model across a broad range of capabilities, ensuring consistent measurement of performance throughout the model lifecycle.
  • Build Automated Benchmarking Pipelines: Design and implement automated evaluation pipelines that execute benchmark tests whenever a new model, fine-tuned checkpoint, or release candidate is available.
  • Maintain Benchmark Quality and Continuous Improvement: Continuously expand, validate, and improve benchmark datasets to ensure they remain representative of real-world user interactions and emerging AI use cases.

Creating Evaluation Programs

  • Human evaluation program: Design and manage structured human evaluation processes to assess model responses for accuracy, helpfulness, safety, cultural relevance, and overall user experience, ensuring continuous improvement through expert and user feedback.

  • Automated evaluation: Design and implement automated evaluation pipelines that continuously assess model performance across predefined benchmarks, safety tests, regression checks, and quality metrics for every training iteration and production release.

  • Safety evaluation: Design and execute comprehensive safety assessments to evaluate the model's robustness against hallucinations, harmful content, bias, prompt injection, jailbreak attacks, privacy risks, and regulatory compliance before production deployment.

  • Prompt evaluation: Design and execute comprehensive prompt-based test suites to systematically evaluate the model's performance across reasoning, instruction following, multilingual understanding, safety, domain-specific tasks, and real-world user scenarios.

Creating a Continuous Improvement framework

  • Analyze Model Performance Trends: Continuously monitor benchmark results, production metrics, and user feedback to identify performance regressions, weaknesses, and improvement opportunities.

  • Drive Evaluation Enhancements: Regularly refine evaluation datasets, benchmarks, prompts, and testing methodologies to ensure comprehensive coverage of emerging use cases and evolving AI capabilities.

  • Provide Actionable Improvement Recommendations: Collaborate with Research, Data, and MLOps teams by translating evaluation insights into prioritized recommendations that improve model quality, safety, and real-world performance.

Technical Skills Required

Large Language Models (LLMs - Llama, Gemma) : Expert

Natural Language Processing (NLP) :Expert

Deep Learning Frameworks (PyTorch, TensorFlow) : Expert

Data Engineering & Pipeline Development : Expert

Python & Data Science Libraries : Expert

Multilingual & Multimodal AI : Advanced

Model Fine-Tuning & Optimization : Advanced

MLOps & Model Deployment : Advanced

Token-based Monetization Strategies : Intermediate

Qualification & Experience

Education

  • Required: Master's degree in Computer Science, AI, Data Science, or a related quantitative field.
  • Preferred: PhD with a research focus on NLP, LLMs, or computational linguistics.

Experience

Required:

  • 10+ years of experience in AI/ML, with at least 7 years in a senior leadership role managing AI research and/or development teams.
  • Proven track record of leading the end-to-end development and successful deployment of large-scale machine learning models, specifically LLMs.
  • Extensive hands-on experience with the entire model lifecycle, including pre-training, fine-tuning, RLHF, and evaluation on massive datasets.
  • Strong background in data engineering, including curating, cleaning, and processing large-scale unstructured text data for model training.
  • Experience in defining and implementing monetization strategies for AI services, such as API-based models or Token as a service.

Preferred:

  • Specific, demonstrable experience in building or adapting language models for non-English languages, ideally Indonesian or related languages.
  • A history of successful collaboration with platform engineering and product teams to integrate complex AI models into production environments.
  • A strong portfolio of relevant publications in top-tier AI conferences (e.g., NeurIPS, ICML, ACL) or significant contributions to major open-source AI projects.

More Info

Job Type:
Industry:
Function:
Employment Type:

About Company

Job ID: 152088445

Beware of Scammers

We don’t charge money for job offers