THE RULES

Rule 16 - Exemption for research, archiving and statistical purposes

Official text

The provisions of the Act shall not apply to the processing of personal data necessary for research, archiving or statistical purposes if it is carried on in accordance with the standards specified in Second Schedule.

Cross-references

Rule 16

Commentary

Rule 16 gives practical effect to Section 17(2)(b) of the Digital Personal Data Protection Act, 2023. It creates a conditional exemption for processing personal data genuinely necessary for research, archiving or statistical purposes. The exemption is available only where the processing complies with the standards in the Second Schedule and the personal data is not used to take a decision specific to a Data Principal.

The Rule should not be read as a broad exemption for universities, research institutions, archives, government agencies, analytics companies or AI developers. It protects a qualifying processing activity, not an organisation merely because it describes itself as a research or archival body. An organisation may conduct exempt research while simultaneously carrying out recruitment, marketing, administration, customer profiling or commercial decision-making to which the ordinary provisions of the DPDPA continue to apply.

Rule 16 is scheduled to come into force on 13 May 2027, eighteen months after publication of the final Rules.

1.1 The nature and limits of the exemption

Rule 16 states that the provisions of the Act shall not apply to qualifying processing. On its face, this is a broad exemption. Where all the conditions are satisfied, the processor of the personal data does not have to apply the ordinary DPDPA framework to that particular activity in the same manner as it would apply to ordinary commercial or operational processing.

That does not mean that the personal data becomes legally unprotected. The exemption is conditional on compliance with the Second Schedule, which preserves a substantial set of data-governance standards. The processing must remain lawful, confined to the exempt purpose, limited to necessary personal data, reasonably accurate and consistent, appropriately secured, retained only while required, and governed by an accountable person determining its purpose and means.

The exemption must therefore be distinguished from a complete legal vacuum. It replaces the ordinary application of the Act with a purpose-specific framework whose conditions continue throughout the life of the research, archive or statistical project. If those conditions cease to be satisfied, the basis for claiming the exemption may also cease.

A genuine research project may qualify when it begins but move outside Rule 16 when its dataset or outputs are later used to rank identified employees, alter insurance premiums, determine access to credit, target advertisements or take another person-specific decision. Similarly, an archive may lawfully preserve historically valuable records but cannot assume that every later commercial reuse remains archival.

1.2 The prohibition against individual-specific decisions

The most important limitation comes from Section 17(2)(b). Rule 16 should not be read independently of that provision. The statutory exemption applies where personal data is processed for research, archiving or statistical purposes and is not used to take a decision specific to a Data Principal.

This separates exempt knowledge-oriented processing from operational processing directed at identifiable individuals.

A decision specific to a Data Principal may include a determination concerning:

  • employment or recruitment;

  • credit or lending;

  • insurance pricing;

  • access to a service or benefit;

  • taxation or enforcement;

  • medical treatment;

  • disciplinary action;

  • account restrictions;

  • personalised pricing;

  • targeted advertising;

  • or another consequence directed at an identifiable person.

The final decision need not be made entirely by a computer. If a statistical model generates a person-level score that a human decision-maker uses to accept, reject, rank or otherwise act upon an identifiable individual, the processing may still be connected with a decision specific to that Data Principal.

Consider an employer that studies historical workforce information to understand overall patterns in employee attrition. It analyses age bands, roles, tenure, location and aggregate departure rates. The findings are used to understand general workforce trends and are not used to evaluate named employees. Subject to necessity and the Second Schedule, this may qualify as research or statistical processing.

If the employer later uses the same dataset to assign every current employee a resignation-risk score and uses that score in promotion, retention or termination decisions, the nature of the processing changes. It is no longer limited to studying an aggregate phenomenon. Personal data is being used to influence decisions specific to identified employees. The Rule 16 exemption would not cover that operational use merely because the model originated in a research project.

The restriction also prevents the exemption from being used indirectly. A research organisation cannot necessarily avoid the limitation by producing identifiable risk scores for another organisation and claiming that only the recipient takes the final decision. If the processing is designed to support decisions about named or identifiable persons, its substance is person-specific.

1.3 Genuine research, archiving and statistical activity

The three purposes recognised by Rule 16 may overlap, but they should be understood according to the real object of the processing.

Research ordinarily involves systematic investigation intended to generate, test or refine knowledge. It may be scientific, medical, technical, historical, social, economic, policy-related or commercial. The Rule does not restrict research to universities, government institutions or non-profit organisations. A private enterprise can conduct genuine research, but applying research methods to personal data does not automatically make the activity exempt.

Example

For example, a hospital may analyse historical treatment outcomes to study whether a particular intervention was generally effective across a patient population. If the analysis is confined to producing general findings and no identifiable patient’s treatment is determined by the research output, the activity may qualify.

If the hospital subsequently integrates the research model into its clinical system and uses it to alter treatment for identified patients, the deployment requires a fresh legal assessment. The model’s history as a research tool does not preserve the exemption once it is used for individual clinical decisions.

Archiving involves deliberate preservation because records possess continuing historical, cultural, evidentiary, institutional or public value. It is not the same as keeping old databases because storage is inexpensive or because the information might become commercially useful.

A public archive may preserve identifiable accounts, photographs or public records concerning an important historical event. Identifiability may be essential to the authenticity and value of the collection. Rule 16 does not require such material to be anonymised if doing so would destroy the archival purpose. It does, however, require the archive to comply with the Second Schedule, including lawful processing, necessity, security, controlled retention and accountability.

By contrast, a business does not create an archival purpose merely by moving an obsolete customer database into low-cost storage and labelling it “archive.” If the real intention is to preserve the information for future marketing, profiling or commercial opportunities, the activity should be assessed according to that actual purpose.

Statistical processing is directed towards identifying patterns, trends, distributions, relationships or aggregate measures. It may support public policy, service planning, epidemiology, economics, transport planning, education or commercial analysis.

A transport authority may analyse pseudonymised journey records to understand route demand and determine where additional public transport capacity is required. The result concerns population movement rather than individual travellers. This may qualify as statistical processing.

If the same journey records are used to determine penalties, adjust individual fares or restrict travel rights for particular passengers, the person-specific use falls outside the exemption.

1.4 Necessity of using personal data

Rule 16 applies only where processing personal data is necessary for the qualifying purpose. The existence of a useful or interesting dataset is not enough.

The organisation should consider whether the objective can reasonably be achieved with:

  • anonymous information;

  • aggregate data;

  • fewer fields;

  • a smaller sample;

  • pseudonymised information;

  • less granular information;

  • or data from which direct identifiers have been removed.

Necessity does not always require complete anonymisation. Some projects genuinely require identifiable or linkable information.

A longitudinal medical study may need to connect records concerning the same patient over several years. An archive may need to preserve the identity of the author of a historically important letter. A statistical authority may need identifiers temporarily to remove duplicate records or connect information from different sources.

In such cases, the organisation should determine whether identifiers are required throughout the project or only at a particular stage. A study may need names temporarily to link records but may no longer require them after a reliable research identifier has been assigned.

1.5 Illustration: Public-health study

Researchers wish to study regional patterns of a particular condition. They may need:

  • age band;

  • district;

  • diagnosis date;

  • relevant clinical characteristics;

  • treatment;

  • and outcome.

They may not need full names, exact residential addresses, identification numbers, bank details or unrelated medical history.

If names are temporarily necessary to link hospital records, the identifying information can be separated after linkage. The research team can then work with coded records while the re-identification key is subject to stronger and more limited control.

The fact that additional personal data could improve convenience or permit future studies does not make it necessary for the present purpose.

1.6 Standards under the Second Schedule

The Second Schedule is the central safeguard supporting the exemption. It requires technical and organisational measures that ensure effective compliance with a set of substantive standards. These standards preserve several of the core disciplines that would otherwise apply under the Act.

The processing must first be lawful. Rule 16 cannot legalise personal data obtained through theft, deception, unauthorised access, breach of confidence or violation of another law. A scientifically valuable project based on unlawfully acquired records does not become legitimate merely because its objective is described as research.

The processing must remain confined to the research, archival or statistical purpose. This requires both governance and operational separation. A dataset collected for a qualifying study should not silently become available to sales, recruitment, insurance or customer-profiling teams.

The information must be limited to what is necessary. An organisation cannot collect every available field and justify it through the possibility that some information may become useful later.

Reasonable efforts must be made to preserve completeness, accuracy and consistency. This is important because poor-quality data can distort research findings, produce false historical accounts and generate misleading statistics. Accuracy in this context does not require historical records to be rewritten. An archive may preserve an original document containing an error but attach appropriate metadata or an explanatory correction rather than altering the original record.

Retention must remain connected to the exempt purpose or another legal requirement. Research may justify preservation during collection, validation, analysis, peer review and a defined reproducibility period. Archiving may justify very long or permanent retention because preservation itself is the continuing purpose. Statistical processing may justify retaining source information while conclusions are validated. None of these concepts supports indefinite preservation without a reasoned connection to the stated purpose.

The Schedule also requires reasonable security safeguards. This is particularly important because research and archival datasets may combine large volumes of information from multiple sources and may contain medical, financial, identity, employment or behavioural information.

Finally, accountability must remain with the person determining the purpose and means of the processing. An organisation claiming the exemption should be able to explain why it applies, what information is necessary, who has access, how the data is protected, what outputs are permitted and when the processing will end.

1.7 Security, access and controlled research environments

Rule 16 does not permit personal data to be distributed freely among researchers, consultants and collaborators merely because they are involved in a qualifying project.

A security model should correspond to the sensitivity, volume and possible consequences of the information. Depending on the project, this may involve separating direct identifiers, using coded research identifiers, encrypting the data, restricting access, controlling exports and reviewing outputs before publication.

Example

For example, a hospital should not send a complete named patient dataset to a research team through an ordinary email attachment merely because the study is legitimate. The project may instead use a secure research environment in which approved users access pseudonymised records, downloads are restricted, activity is logged and aggregate outputs are checked before release.

Pseudonymisation is often useful, but it should not be confused with anonymisation. Replacing names with reference numbers does not remove the data from the category of personal data if the records can be connected back to the individual through a separate key or other reasonably available information.

Similarly, removing obvious identifiers may not create anonymous data where remaining attributes can single out a person. A dataset containing an exact age, small village, rare diagnosis, occupation and treatment date may identify the individual even if the name has been removed.

The test is practical identifiability, not the label applied by the research team.

1.8 Publication and disclosure of findings

Research and statistical outputs should be designed to prevent unnecessary identification.

Aggregate publication is not automatically safe. A category may be so small or distinctive that the individuals within it can be identified.

Suppose a study reports that the only employee in a particular office performing a specialised role has a stated health condition. The result may effectively identify that employee even if her name is omitted.

Medical and social-science case studies create similar concerns. A report may remove the subject’s name but include an exact age, profession, location, rare condition and detailed event history. Family members, colleagues or members of the local community may identify the person from the combination.

Archival preservation and public access should also be distinguished. An archive may lawfully preserve a record while restricting access for a specified period or permitting use only under controlled conditions. The existence of enduring archival value does not necessarily require immediate online publication of every identifiable record.

1.9 Reuse and change of purpose

The Rule 16 assessment must continue when the project changes. An exempt dataset does not carry permanent exempt status into every later use.

A new assessment is required where:

  • a different research question is introduced;

  • new categories of personal data are added;

  • another institution receives the information;

  • direct identifiers are reintroduced;

  • individual scores are generated;

  • outputs are used operationally;

  • data is supplied for advertising or profiling;

  • a research model is commercialised;

  • or retention is materially extended.

The dividing line is especially important in AI development.

A university may train an experimental model to study whether large language models memorise personal data. The work is conducted in a secure research environment, the model is not used to decide anything about individuals, and outputs are controlled. That project may potentially qualify.

If the model is then licensed to a company that uses it to evaluate identifiable job applicants, customers or patients, the operational deployment falls outside the research exemption. The new user must establish an ordinary lawful basis, issue applicable notices, support rights, control Processors, address retention and comply with the rest of the DPDPA.

The same principle applies where a company develops a recommendation model and calls the development “data science research.” If the system is designed to determine which price, advertisement or offer a named customer receives, the processing is connected with individual-specific decisions. Rule 16 should not apply merely because statistical or machine-learning techniques were used.

1.10 Research collaborations and Data Processors

Qualifying projects often involve several organisations. Each participant’s role must be determined from the actual arrangement.

A cloud provider storing research information only on documented instructions may act as a Data Processor. A university and hospital jointly determining the research question, dataset and methodology may each exercise decision-making over the purpose and means. A sponsor receiving only aggregate results may stand in a different position from one controlling participant selection and research outputs.

The exemption should be assessed for each person determining the relevant processing. One organisation cannot automatically inherit another organisation’s exemption simply because the information originated in a qualifying project.

The arrangements should address:

  • the approved purpose;

  • permitted data;

  • access;

  • security;

  • onward sharing;

  • publication;

  • individual-decision restrictions;

  • retention;

  • breach response;

  • and deletion or return at the end of the project.

A Processor should not independently use the dataset to train its own commercial model or develop unrelated products unless that new processing has its own lawful basis and, where Rule 16 is claimed, independently satisfies its conditions.

Where Rule 16 validly applies, the DPDPA does not require the research, archival or statistical activity to rely on ordinary consent under the Act.

However, this does not remove duties arising under other laws, professional obligations or ethical frameworks. Medical research, clinical trials, government records, banking information, official secrets, professional confidentiality and institutional research may remain subject to separate requirements.

A clinical study may still require informed participant consent under health-research standards even if the processing satisfies the DPDPA research exemption. Conversely, obtaining research consent does not make an individual-decision project exempt under Rule 16.

The two questions are different:

  • whether another framework requires consent to conduct the study; and

  • whether the DPDPA research exemption applies to the personal-data processing.

Rule 16 cannot override another applicable law requiring a higher level of confidentiality, participant approval or institutional review.

1.12 Data Principal rights and mixed datasets

Because Rule 16 states that the Act does not apply to qualifying processing, ordinary rights under the DPDPA may not operate against the exempt research, archival or statistical activity in the same manner as they apply to ordinary processing.

This makes careful classification essential. An organisation should not refuse a rights request simply by describing a database as “research.”

The same information may exist in both exempt and non-exempt systems.

A hospital may hold:

  • a patient’s active clinical record used for treatment; and

  • a pseudonymised copy used in a qualifying population-level study.

The research copy may potentially qualify under Rule 16. The active treatment record does not become exempt simply because some of its information was also supplied to a research project.

Similarly, an employer may use workforce information for aggregate statistical analysis while separately using the same information for payroll, appraisal and disciplinary decisions. Those operational systems remain subject to the ordinary DPDPA framework.

When responding to a Data Principal, the organisation should distinguish the copies and purposes accurately rather than applying a project label to the person’s entire information history.

1.13 Governance and evidence of eligibility

A person relying on Rule 16 should be able to demonstrate why the exemption applies.

The project record should explain:

  • the genuine research, archival or statistical objective;

  • why personal data is necessary;

  • why anonymous or less identifiable information is insufficient;

  • what personal data is used;

  • where it came from;

  • which organisations and Processors are involved;

  • how individual-specific decisions are prevented;

  • how security and access are managed;

  • how long the information will be retained;

  • what outputs may be released;

  • and what will happen if the project later becomes operational or commercial.

An ethics committee or institutional review board may provide useful oversight, particularly for health, behavioural or social-science research. Its approval is relevant evidence but does not automatically establish the Rule 16 exemption. The project must still satisfy the statutory no-individual-decision condition and the Second Schedule.

Eligibility should be reassessed when material changes occur. A research approval completed at the beginning does not continue to protect a project that has evolved into customer profiling, personalised decision-making or unrelated commercial development.

1.14 Practical application

A few integrated examples illustrate the legal boundary.

1.15 Population-level medical research

A hospital studies whether a treatment is associated with better outcomes across different age groups. Identifiers are separated, researchers use coded records and findings are published only in aggregate form. No patient’s present treatment or insurance position is affected.

This may qualify if the personal data is necessary and the Second Schedule is followed.

If the hospital later returns named risk scores to treating doctors and those scores determine treatment, the operational activity requires a separate legal basis.

1.16 Historical archive

An archive preserves personal letters, photographs and records concerning a major public event. Identities are necessary to maintain historical authenticity. Access to particularly consequential records concerning living individuals is controlled.

The activity may qualify as archiving.

Using the same records to build present-day employment, credit or political profiles would be a new and non-archival purpose.

1.17 Aggregate workforce statistics

An employer studies broad patterns of attrition by role, location and tenure. The analysis is used for general workforce planning and not to rank employees.

The project may qualify as statistical processing.

If the employer assigns each employee an attrition score and uses it to determine promotion or termination, the exemption is unavailable for that use.

1.18 Customer analytics

A retailer analyses total product demand by city and season to plan warehouse capacity. The output concerns aggregate demand.

This may potentially be statistical processing.

If the retailer uses the records to vary prices or target advertisements for each identified customer, the processing becomes individualised commercial decision-making.

1.19 Loss of exemption

The exemption can be lost where the processing:

  • is not genuinely for research, archiving or statistics;

  • uses more personal data than necessary;

  • becomes connected with individual-specific decisions;

  • is sourced or conducted unlawfully;

  • is repurposed for profiling, advertising or operational action;

  • fails to maintain reasonable security;

  • retains personal data without continuing justification;

  • produces unnecessarily identifiable outputs;

  • or otherwise fails the Second Schedule.

Loss of the exemption does not necessarily mean the organisation may continue processing by retrospectively describing another legal ground. It must determine whether an ordinary DPDPA basis actually exists and whether the processing can satisfy notice, consent or legitimate-use conditions, rights, security, erasure and other obligations.

In some cases, a dataset assembled for exempt research may not lawfully be converted into an operational system at all without redesign, new data collection or a fresh legal basis.

1.20 Enforcement implications

Improper reliance on Rule 16 may expose the organisation to the provisions it wrongly treated as inapplicable.

Possible contraventions could involve:

  • processing without consent or another applicable ground;

  • failure to provide notice;

  • denial of Data Principal rights;

  • excessive or indefinite retention;

  • inadequate security;

  • failure to notify a breach;

  • unlawful Processor use;

  • or impermissible processing involving children.

The applicable penalty would depend on the underlying contravention. A security failure, breach-notification failure or children’s-data violation may attract the specific penalty category applicable to that obligation. Other significant contraventions may fall within the residual penalty category.

Deliberately describing personalised commercial profiling as “research” would be materially different from a reasonable but mistaken interpretation in a genuinely structured research project. Relevant factors would include the actual purpose, scale, personal-data categories, consequences for individuals, safeguards, internal documentation, remedial action and whether the organisation continued after recognising the problem.

1.21 Concluding interpretation

Rule 16 preserves an important legal space for creating knowledge, producing reliable statistics and maintaining records of enduring value. At the same time, it prevents the research label from becoming an escape route from data-protection law.

The exemption applies only where:

  • the processing is genuinely directed at research, archiving or statistics;

  • using personal data is necessary;

  • the information is not used to make decisions specific to identifiable Data Principals;

  • and the processing continuously satisfies the Second Schedule.

An organization relying on the Rule should be able to explain the project’s purpose, the need for personal data, the prohibition on individual use, the security model, the retention period, the permitted outputs and the controls against later commercial or operational reuse.

Key point

Rule 16 protects genuine research, archival preservation and statistical analysis, but not individual profiling disguised in technical language. The decisive question is not whether statistical or research methods are used. It is whether personal data is genuinely necessary for the exempt purpose, remains subject to the Second Schedule and is kept outside decisions concerning identifiable individuals. Once the data or output is used to act upon a particular Data Principal, the ordinary protections of the DPDPA apply to that use.

Reproduced from official sources for reference. Not legal advice. In case of any discrepancy, the text published in the Gazette of India prevails.