About Slash
Slash is a hi-tech startup studio with a mission to build tech AI-powered products and scalable digital platforms that create real-world impact. Since 2016, we've partnered with ambitious enterprises and government organizations to design, engineer, and launch cutting-edge solutions — with Generative AI at the core of what we do.
We specialize in AI-powered application delivery, from product design and high-performance engineering to DevOps and AI operations. Headquartered in Singapore, our global team and clients operate R&D hubs across Southeast Asia. We are a team of entrepreneurs, engineers, and product builders dedicated to solving complex technical challenges and turning bold ideas into impactful technology.
About our Client
Our client's product is an AI voice-sensing device, a breakthrough wearable that detects the gap between what someone says and how their voice actually sounds. Rooted in Pythagorean acoustic physics and the Navarasa framework, the system functions as a state detector rather than a conventional emotion labeler. The team is a small, fast-moving team building at the intersection of emotion labeling, ancient wisdom, and measurable science.
About the Role
We are looking for a hands-on Machine Learning Engineer to take complete ownership of our full AI stack. Your primary responsibility will be expanding our speech emotion recognition (SER) model into a physics-based harmonic vocal state engine, alongside building and maintaining our dual-instance production LLM infrastructure. In this role, you will work directly with the Founder and Lead Developer with zero bureaucracy or committee oversight. We need an engineer who excels at owning problems end-to-end.
Key Responsibilities
1. Harmonic Vocal State Engine (Audio & Physics)
- Extend our inherited 30-class speech emotion recognition (SER) model, dataset, checkpoints, and pipeline into a physics-based harmonic detection system using Fourier-derived acoustic analysis and the Navarasa framework.
- Design and implement an in-house model validation methodology from scratch using approaches like Gemini's emotion labeling API, cross-validation on open datasets (IEMOCAP, RAVDESS), or custom ground-truth evaluation pipelines.
2. LLM Infrastructure & Operations
- Deploy and manage a dual-instance vLLM setup on GCP (g2-standard-24 instance): GPU 0: Llama 3.1 8B for fast-lane prompts; GPU 1: Qwen 2.5 32B for reasoning-heavy prompts.
- Own prompt engineering, output validation, Pydantic schema enforcement, retry logic, and quality monitoring across 42 production prompts.
- Handle infrastructure scaling and migrations independently without reliance on third-party vendors.
3. Model Quality & Continuous Improvement
- Build evaluation pipelines to detect and catch regressions before users experience them.
- Continuously optimize and refine model performance as real-world audio accumulates from our device.
Requirements & Qualifications
- Audio Signal Processing: Strong proficiency in the frequency domain (FFT, Fourier analysis, spectral features) and practical experience with audio processing libraries like librosa or torchaudio.
- Audio ML Frameworks: Hands-on experience with modern audio ML models (wav2vec2, HuBERT, WavLM) and a deep understanding of training, evaluating, and fine-tuning SER models.
Core Tech Stack:
- PyTorch & Python: Direct experience handling model checkpoints and training loops; ability to write clean, modular, maintainable production-level code.
Infrastructure & MLOps:
- GCP & vLLM: Experience managing cloud infrastructure and direct experience deploying and serving models via vLLM.
- Methodology & Communication: Proven ability to design rigorous evaluation frameworks and strong written English skills for an async-first team.
Nice to Have
- Experience designing or building physics-based audio analysis systems.
- Familiarity with the Navarasa framework, acoustic physics, or classical music/acoustic theory systems.
- Experience using FastAPI for inference API layers.
- Prior experience replacing or migrating away from third-party AI benchmarks or providers.
- Genuine personal interest in sound, acoustic physics, consciousness, or the intersection of emotions and modern ML.
Hiring Process (Take-Home Assessment)
We do not conduct whiteboard interviews. Our technical evaluation is a 48-Hour Practical Assessment consisting of two deliverables:
- SER Model Assessment: Review our inherited SER validation report and dataset documentation, providing your strategy for validation.
- Architecture Assessment: Provide a 1-page technical evaluation of a Python file from our harmonic engine, detailing its architecture and roadmap.