No title in Indian tech is more overloaded than data scientist. It is attached to people who build dashboards, people who run experimentation for a product team, people who ship models to production, and people who publish research. Those are four different jobs. Most bad data hires are not assessment failures, they are definition failures.
The four jobs behind one title
- Analytics and reporting. Answers business questions with existing data. Strong SQL, visualisation, business fluency. Output is decisions made faster. Often mislabelled as data science because the title attracts better candidates.
- Product data science. Designs experiments, measures whether changes worked, handles causal questions. Heavy statistics, close partnership with product managers. The most commonly needed and least commonly specified.
- Machine learning engineering. Ships models to production and owns them there: pipelines, serving, latency, drift, monitoring. Closer to software engineering than to statistics.
- Research. Novel modelling, publications, deep specialisation. Genuinely needed by few companies, wanted by many, and expensive to hire wrongly.
A candidate can be excellent at one and unable to do another. The analytics specialist who has never deployed anything is not a bad engineer, they simply have not done that job. Hiring them for production ML is your mistake, not theirs.
Decide by output, not by skill list
Skip the technology list entirely for a moment and answer one question: twelve months from now, what exists because this person was here? A set of dashboards the leadership team uses weekly. A running experimentation practice with results nobody disputes. A recommender serving live traffic with monitored performance. A published model that advanced the state of the art.
Each answer implies a completely different hire, a different band, and a different sourcing pool. Write that sentence before writing the job description, as covered in outcome based hiring.
The India-specific screening problem
Data science attracted enormous course and certification volume in India, which produced a large population of candidates with strong resumes and thin applied experience. Keyword screening cannot distinguish them, because a certificate and three years of production work both put the same tool names on a profile.
What distinguishes them is evidence of consequence. Did a model reach users. Was a decision made differently because of their analysis. Did they own something after launch. One well-chosen question, asking them to walk through a project end to end including what went wrong, separates the two populations faster than any test.
Assessment that respects their time
The Indian data hiring norm of a multi-hour open-ended take-home filters for availability, not ability, and the strongest candidates with competing processes simply decline. Three better options.
- Project deep dive. Forty minutes on something they built. Why that approach, what was rejected, how they knew it worked, what they would change.
- Live problem framing. Give a messy business question and watch them turn it into something answerable. This is the actual daily job.
- Short bounded exercise. Ninety minutes maximum, with a clear scope, on data resembling yours.
Where they actually are
Filter by the kind of company rather than by keywords. Someone who worked on ranking at a consumer company with genuine scale has faced problems that do not appear at a firm with 50,000 rows, regardless of matching tool names on both profiles. Company context is the most useful filter in data hiring and the least used.
Strong applied data people are usually employed and content, so expect direct sourcing rather than applications, and expect the 30 to 90 day notice period that comes with it.
Search by what they built, not what they listed
Describe the data problem in plain English and get candidates scored on whether their history proves they can do it, with the reasoning shown. Two free searches, no card.
Start a free search →