AI/ML Engineer
Posted June 29, 2026
About Us
We are a fast-growing startup based in Pune, India, specializing in cutting-edge Data Science and Data Engineering solutions. Our team of dedicated professionals is committed to solving complex data challenges for companies worldwide.
Our Culture
We foster a vibrant startup culture that values:
- Intellectual curiosity
- Continuous learning
- Positive work environment
- Collaborative problem-solving
Role Overview
We are seeking a highly technical and proactive AI-ML Engineer to join our dynamic team. The ideal candidate will possess a strong blend of software engineering rigor and deep technical expertise in modern AI/ML architectures, infrastructure optimization, and generative AI systems. This role demands critical thinking, production-grade coding, and the ability to deploy, optimize, and scale robust AI models to solve complex, real-world problems. You should bring a deep curiosity for understanding model internals, and the adaptability required to navigate and implement rapidly evolving, state-of-the-art technologies.
Key Responsibilities
- Deliver end-to-end AI engineering projects by designing, building, and deploying production-grade Machine Learning, Deep Learning, and Generative AI applications.
- Develop high-quality software solutions in Python, collaborating with cross-functional engineering teams to integrate AI models into existing application codebases.
- Implement advanced training strategies including mixed-precision training (FP16/BF16), gradient accumulation, and distributed training (Data/Model/Pipeline parallel) while profiling and maximizing GPU utilization.
- Optimize, serialize (ONNX, TorchScript), and deploy models via high-throughput REST APIs (FastAPI) while managing latency vs. throughput trade-offs.
- Implement robust MLOps workflows using DVC, Docker, and cloud platforms for model versioning, pipeline automation, and production monitoring.
- Architect and optimize high-performance Retrieval-Augmented Generation (RAG) systems using hybrid search, semantic chunking, and cross-encoder reranking.
- Design autonomous, multi-agent systems and orchestrated workflows utilizing function calling, the ReAct pattern, and advanced memory management architectures.
- Implement state-of-the-art serving techniques (vLLM, speculative decoding, prompt caching) and quantization (INT8/INT4 via GPTQ/AWQ) to maximize inference throughput.
- Move workflows from notebooks to production-grade pipelines, writing clean code and implementing unit/integration tests for ML (pytest, Great Expectations).
- Actively diagnose production anomalies, including data/concept distribution shifts, silent model failures, and training bottlenecks (underfitting/overfitting).
- Evaluate and benchmark LLM outputs using appropriate metrics and testing frameworks.
- Design high-throughput data pipelines and optimized SQL/NoSQL queries for large-scale data processing and model feature injection.
- Practice active listening to understand project requirements and team inputs.
- Collaborate with stakeholders to translate complex business requirements into scalable AI/ML solutions and communicate technical trade-offs clearly.
- Demonstrate strong technical communication skills, a high degree of ownership, and an action-biased approach to debugging and solving ambiguous engineering problems
- Apply responsible AI principles, jailbreak awareness, and output validation guardrails to ensure ethical and safe model development.
- Plan strategically and multitask efficiently to meet project deadlines.
Required Skills
Core Programming & ML
- Strong Python programming skills with hands-on project experience, Git proficiency, and a basic understanding of CUDA
- Expertise in Deep Learning architectures (Transformers, CNNs, RNNs) alongside a strong theoretical understanding of foundational ML algorithms (GBMs, Random Forests)
- Deep proficiency in PyTorch or TensorFlow, with a strong emphasis on custom layer implementation and neural network training loops
- Hands-on experience with modern training paradigms, including Self-supervised Learning, Contrastive Learning, and advanced Transfer Learning
- Experience with NLP, Computer Vision, or Time Series Analysis
- Proven experience writing clean, production-grade code, utilizing pytest and Great Expectations for data and model validation
Generative AI & LLMs
- Hands-on experience with commercial LLM APIs (OpenAI, Anthropic, Groq) and hosting/deploying open-source foundational models (Llama, Mistral)
- Proficiency with modern orchestration frameworks (LangChain, LlamaIndex) and building stateful multi-agent architectures using LangGraph or DSPy
- Experience with vector databases (Pinecone, Weaviate, Chroma, pgvector), advanced chunking (fixed, semantic, recursive), and hybrid search implementation
- Deep technical understanding of Parameter-Efficient Fine-Tuning (PEFT) mechanics—specifically LoRA/QLoRA low-rank decomposition—alongside instruction tuning and model alignment methodologies (DPO, RLHF)
- Mastery of advanced prompt engineering, including structured output forcing, Chain-of-Thought, system prompt design, and building resilient multi-agent coordination systems
- Hands-on experience with inference optimization and high-throughput serving, including quantization (GPTQ, AWQ), speculative decoding, prompt caching, and vLLM acceleration
MLOps & Deployment
- Experience with MLOps practices, logging, and model registries (MLflow, Weights & Biases, DVC) along with model serving via Triton Inference Server or FastAPI
- Production experience deploying, scaling, and monitoring models natively on cloud AI platforms, specifically AWS SageMaker or GCP Vertex AI
- Experience building CI/CD pipelines for ML applications, with an emphasis on data and model versioning/lineage using DVC and Git
Data Engineering & Databases
- Solid understanding of SQL, including advanced concepts like windowing functions and query optimization
- Experience building and orchestrating data pipelines using Airflow or Prefect to feed specialized SQL and NoSQL vector/feature databases
Soft Skills & Professional Attributes
- Strong critical thinking and problem-solving skills
- Excellent written and verbal communication abilities
- High degree of flexibility and adaptability to stay ahead of the rapid velocity of the open-source AI ecosystem (Hugging Face, vLLM, etc.)
- Practical understanding of AI safety, compliance, guardrails (e.g., NeMo Guardrails or Llama Guard), and responsible AI practices.
Nice-to-Have
- Contributions to open-source AI/ML repositories or GenAI orchestration frameworks.
- Experience with real-time streaming data processing
- Active participation in competitive ML spaces (Kaggle) or track record of reviewing/reproducing state-of-the-art AI research papers
- Published research papers or conference presentations
- Experience building and scaling Graph Databases or Knowledge Graphs for advanced RAG
- Experience with multimodal architectures (Vision-Language models, Audio processing)
Qualifications
- AI-ML Engineer: 2–5 years of hands-on experience engineering machine learning systems, optimizing infrastructure, and implementing LLM/GenAI workflows in production
- Bachelor's or Master's degree in Computer Science, Artificial Intelligence, Data Science, Statistics, or a related highly quantitative field
- Demonstrated commitment to continuous learning through contributing to open-source, certifications, or self-study (especially in deep learning internals, MLOps, and modern GenAI frameworks)
What We Offer
- Competitive salary commensurate with experience
- Opportunity to work on diverse, cutting-edge AI/ML projects
- Collaborative and innovation-driven work environment
- Rapid growth and continuous learning opportunities
- Exposure to latest AI technologies and industry best practices

