JOB RESPONSIBILITIES
- Conduct research and experimentation on state-of-the-art LLM architectures, data curation techniques, and training recipes.
- Develop and implement novel approaches across the model lifecycle: pre-training, post-training (preference optimization, RLHF), prompt engineering, retrieval-augmented generation (RAG), agentic workflows, and memory engineering.
- Apply cutting-edge NLP advancements to improve Khmer text understanding, including tokenization, word segmentation, and interpretability of model behavior on Khmer.
- Identify, collect, and curate high-quality, diverse datasets in Khmer for LLM training, validation, and testing.
- Build and maintain scalable data pipelines for large-scale collection, cleaning, deduplication, and filtering of Khmer text.
- Generate and curate synthetic and augmented Khmer data to address gaps in low-resource domains.
- Train and fine-tune Large Language Models using extensive Khmer datasets.
- Optimize LLM performance for tasks such as text generation, summarization, translation, question answering, and sentiment analysis in Khmer.
- Collaborate closely with MLOps Engineers to integrate trained LLMs.
- Work on containerization (Docker) and orchestration (Kubernetes) of LLM services.
- Build and maintain data visualization and analytics pipelines to monitor model performance, training metrics, and data quality using enterprise big data tools.
- Design and implement rigorous evaluation methodologies and metrics tailored for Khmer LLMs.
- Conduct comprehensive testing and analysis to identify model biases, ethical concerns, and areas for improvement.
JOB REQUIREMENTS
- Bachelor's or Master's degree in Computer Science, Artificial Intelligence, Computational Linguistics, or a related field.
- Solid understanding of LLM architectures and their underlying mechanisms.
- Familiarity with MLOps principles, CI/CD pipelines, Docker, and Kubernetes.
- Understanding of Khmer linguistics and script characteristics, including word segmentation and tokenization challenges.
- Knowledge of data structures, algorithms, and software engineering best practices.
- Minimum of 3+ years of hands-on experience in Natural Language Processing (NLP) or Machine Learning engineering.
- Proven experience working with Large Language Models (LLMs), including pre-training, fine-tuning, or deployment.
- Strong proficiency in Python and relevant NLP/ML libraries
- Experience with cloud platforms for training and deploying ML models.
- Experience building and managing large-scale datasets and data pipelines.
- Strong analytical and problem-solving mindset, with attention to detail in experimentation and evaluation.
- Clear communicator, able to explain complex technical concepts to both technical and non-technical audiences.
- High integrity in handling data, respecting privacy, licensing, and ethical considerations.
- Intellectual curiosity and self-motivation to keep pace with a fast-moving research field.
- Collaborative team player, able to work closely across teams and with stakeholders.