Key Responsibilities
- Define Model Evaluation Strategy
- Design a Comprehensive Evaluation Framework: Develop and maintain a standardized evaluation framework that measures model performance across every stage of the LLM lifecycle, from pre-training to production deployment.
- Define Quality Metrics and Success Criteria: Establish clear Key Performance Indicators (KPIs) and acceptance thresholds that determine whether a model is ready to progress to the next development stage or production deployment.
- Develop Indonesia-Centric Evaluation Benchmarks: Create and maintain benchmark datasets that accurately measure the model's understanding of Bahasa Indonesia, regional dialects, cultural context, local regulations, and industry-specific terminology.
- Define Quality Gates: Define the evaluation process that every model must successfully complete before being approved for deployment.
Benchmark Development
- Design Comprehensive Benchmark Suites: Develop and maintain standardized benchmark datasets that evaluate the model across a broad range of capabilities, ensuring consistent measurement of performance throughout the model lifecycle.
- Build Automated Benchmarking Pipelines: Design and implement automated evaluation pipelines that execute benchmark tests whenever a new model, fine-tuned checkpoint, or release candidate is available.
- Maintain Benchmark Quality and Continuous Improvement: Continuously expand, validate, and improve benchmark datasets to ensure they remain representative of real-world user interactions and emerging AI use cases.
Creating Evaluation Programs
- Human evaluation program: Design and manage structured human evaluation processes to assess model responses for accuracy, helpfulness, safety, cultural relevance, and overall user experience, ensuring continuous improvement through expert and user feedback.
- Automated evaluation: Design and implement automated evaluation pipelines that continuously assess model performance across predefined benchmarks, safety tests, regression checks, and quality metrics for every training iteration and production release.
- Safety evaluation: Design and execute comprehensive safety assessments to evaluate the model's robustness against hallucinations, harmful content, bias, prompt injection, jailbreak attacks, privacy risks, and regulatory compliance before production deployment.
- Prompt evaluation: Design and execute comprehensive prompt-based test suites to systematically evaluate the model's performance across reasoning, instruction following, multilingual understanding, safety, domain-specific tasks, and real-world user scenarios.
Creating a Continuous Improvement framework
- Analyze Model Performance Trends: Continuously monitor benchmark results, production metrics, and user feedback to identify performance regressions, weaknesses, and improvement opportunities.
- Drive Evaluation Enhancements: Regularly refine evaluation datasets, benchmarks, prompts, and testing methodologies to ensure comprehensive coverage of emerging use cases and evolving AI capabilities.
- Provide Actionable Improvement Recommendations: Collaborate with Research, Data, and MLOps teams by translating evaluation insights into prioritized recommendations that improve model quality, safety, and real-world performance.
Technical Skills Required
Large Language Models (LLMs - Llama, Gemma) : Expert
Natural Language Processing (NLP) :Expert
Deep Learning Frameworks (PyTorch, TensorFlow) : Expert
Data Engineering & Pipeline Development : Expert
Python & Data Science Libraries : Expert
Multilingual & Multimodal AI : Advanced
Model Fine-Tuning & Optimization : Advanced
MLOps & Model Deployment : Advanced
Token-based Monetization Strategies : Intermediate
Qualification & Experience
Education
- Required: Master's degree in Computer Science, AI, Data Science, or a related quantitative field.
- Preferred: PhD with a research focus on NLP, LLMs, or computational linguistics.
Experience
Required:
- 10+ years of experience in AI/ML, with at least 7 years in a senior leadership role managing AI research and/or development teams.
- Proven track record of leading the end-to-end development and successful deployment of large-scale machine learning models, specifically LLMs.
- Extensive hands-on experience with the entire model lifecycle, including pre-training, fine-tuning, RLHF, and evaluation on massive datasets.
- Strong background in data engineering, including curating, cleaning, and processing large-scale unstructured text data for model training.
- Experience in defining and implementing monetization strategies for AI services, such as API-based models or Token as a service.
Preferred:
- Specific, demonstrable experience in building or adapting language models for non-English languages, ideally Indonesian or related languages.
- A history of successful collaboration with platform engineering and product teams to integrate complex AI models into production environments.
- A strong portfolio of relevant publications in top-tier AI conferences (e.g., NeurIPS, ICML, ACL) or significant contributions to major open-source AI projects.