Skip to content
Zentis AI
Request a demo
Zentis AI
Industries
Banking & creditSOUL co-pilot, retail credit, onboarding.InsuranceFNOL, claims triage, underwriting.Finance, risk & voiceForecasting, reserving, voice agents.Insurance BrokerageEvery quote and every renewal, ready before the client starts wondering.
Products
Zentis AnalyticsForecasting, reserving, and regulatory reportingAuditOSBank, e-commerce, Shariah, and claims audit.FAMIend-to-end motor insurance automationZentis BinderDefence file assembly, live as the case runs.
Resources
NewsProduct news and conference write-ups.ArticlesLonger pieces on AI in regulated industry.Events & webinarsWhere to find us in person.UsecasesA library of reusable, editable templates.
Company
AboutWho we are and why the harness is the product.PartnersTechnology, consulting, and reseller partners.ContactTalk to the team, or book a working session.CareersJoin us
Platform
ZaraZara is where a workflow starts, Zara builds the team.Zen StudioZen Studiois where engineering opens it upZen PilotZen Pilot is where it actually runs
Request a demo

Stay Ahead with Zentis AI Insights

Subscribe to receive the latest updates straight to your inbox

Zentis AI

An enterprise-grade agentic platform for banking, insurance, and audit. Incubated by Techvantage.ai.

Platform

ZaraSPARZen StudioZen Pilot

Solutions

Audit & complianceInsuranceBanking & creditFinance, risk & voice

Trust

The defence fileSecurityAssuranceRegulatory packs

Resources

NewsArticlesEvents & webinarsClassic

Company

AboutPartnersContactCareers
© 2026 Zentis AI. All rights reserved.info@zentis.aiLondon, England
Anthropic Partner NetworkNVIDIA InceptionTechvantage.ai, Deloitte Technology Fast 50

Careers

Applied ML Engineer, Small Models

Fine-tune domain models on BFSI data with QLoRA, run the evaluation harness, and get inference running inside customer perimeters.

Objective

Join our Model Layer team to fine-tune compact, domain-specialized models for BFSI use cases. You'll own the full ML lifecycle — data curation, QLoRA fine-tuning, evaluation, and deploying inference inside customer perimeters (VPC / on-prem).

About the role

Join our Model Layer team to fine-tune compact, domain-specialized models for BFSI use cases. You'll own the full ML lifecycle — data curation, QLoRA fine-tuning, evaluation, and deploying inference inside customer perimeters (VPC / on-prem).

Most BFSI work is narrow and well-specified — exactly what a small, fine-tuned model is good at. Your job is to make 1B–13B models carry real regulatory work: field extraction, classification, routing, and domain validation, running inside the customer's own perimeter with no cloud to fall back on.

You own the whole loop — data, training, evaluation, serving — and you document every checkpoint so a model shipped in March can be defended in September. The registry is not a suggestion; it is the product.

Key responsibilities

  • Fine-tune small language models (1B–13B) using LoRA / QLoRA / PEFT techniques
  • Curate, clean, and version high-quality BFSI training datasets
  • Build and maintain rigorous evaluation harnesses for domain accuracy
  • Optimize models for inference (quantization, distillation, pruning)
  • Deploy and support models running inside customer environments (air-gapped / VPC)
  • Partner with Solutions Architects to align model behavior with customer requirements

Required skills

  • 2–5 years of experience in Applied ML / NLP
  • Strong Python, PyTorch, and Hugging Face (Transformers, PEFT, TRL, Datasets)
  • Hands-on with QLoRA, LoRA, SFT, and DPO/RLHF
  • Experience with model quantization (bitsandbytes, GPTQ, AWQ, GGUF)
  • Familiarity with inference servers (vLLM, TGI, Ollama, Triton)
  • Solid understanding of evaluation methodology (benchmarks, LLM-as-judge, human eval)
  • Experience with GPU workflows (CUDA, distributed training)

Nice to have

  • BFSI domain knowledge (banking, insurance, capital markets)
  • Experience deploying models in regulated / on-prem environments
  • Published work, Kaggle rank, or OSS contributions in NLP/LLMs

Skills & technologies

PythonPyTorchHugging FaceQLoRA / LoRA / PEFTSFT / DPOQuantizationvLLM / TGI / TritonEvaluation methodologyCUDA
Apply for this roleBack to all roles

Role at a glance

Location
Trivandrum, India
Team
Model layer
Type
Full-time
Apply for this role

PDF, Word or text resume, up to 10 MB.

Rather talk than apply?

Tell us what you would build and why it matters. If the timing is right, the conversation goes from there.

Send a note