DPDP Act 2023 · Data discovery
DPDP compliance begins at the column level.
The Act has been law since 2023, the Rules were notified in November 2025, and the obligations that actually bite — notice and consent, data-principal rights, reasonable security safeguards and a 72-hour breach report — commence in May 2027. Every one of them assumes you already know which tables hold personal data, which accounts can read those tables, and who has looked. Most organisations do not, and that inventory is not something you can produce in the week somebody asks for it.
when notice, consent, data-principal rights and breach reporting commence
maximum penalty where a breach traces back to absent security safeguards
reveal of flagged personal data recorded against a named account
Why DPDP stalls before it starts.
The Act asks you to know what personal data you hold, to protect it, and to answer for it. The knowing is where it stalls, because nobody can protect a column they have not found — and every later obligation quietly assumes that inventory already exists.
You cannot inventory what you have never opened
A Data Fiduciary is accountable for personal data wherever it sits. In practice it sits in databases nobody has audited column by column — a copy taken for a migration, a reporting replica, a staging catalog restored from live and never cleaned. A questionnaire sent to application owners returns what they remember, not what is there.
Column names lie in both directions
A column called cust_ref holding Aadhaar numbers is invisible to a scan that reads names. A column called pan_no holding an internal part number is a false positive that costs a week. An inventory built from the schema is wrong in both directions at once, and nothing in it tells you which entries are which.
The clock starts when you notice, not when you are ready
A reportable breach means telling the Board without delay and filing a detailed report within 72 hours — what data, whose, how many, and what you have done about it. If the inventory does not already exist, those hours go on establishing the blast radius instead of containing it. The 72 hours run continuously; weekends included.
"Who accessed it" is a question about the past
Accountability sits on the Data Fiduciary, and the honest answer to "who looked at whose personal data" is either a log you were already keeping or nothing at all. Access records cannot be reconstructed afterwards. As with any evidence question, that window is simply gone.
What the console does about it.
Not a policy tracker with a questionnaire attached. Discovery that runs against the live estate on a schedule, and the record it leaves behind as the deliverable. The panels are named after the questions they answer.
What we found — and how it was verified
Personal data grouped as it is regulated: statutory identifiers — Aadhaar, PAN, passport, voter ID, driving licence, GSTIN — alongside financial data and personal details. Each tile states how the finding was reached, because several identifiers are confirmed by running the number's own checksum rather than by matching a column name, and the tile tells you which.
Where it is — down to the column
The exact server, catalog, table and column behind every finding, with an Inspect button beside it. Not "personal data exists somewhere in billing" but the address of the column holding it, which is the only form of the answer a remediation ticket can be written against.
Who can read it
Which accounts can reach the personal data, and how they came by that access — directly, or through a role that grants more than anyone intended. The entitlement explorer sits alongside it, calling out over-privileged accounts, accounts with no password at all, dormant accounts and separation-of-duties conflicts as counts you can act on.
Who looked — masked by default, reveal on the record
Inspect returns a masked sample: enough to judge whether a finding is real, without exposing the data to the person triaging it. Unmasking is a separate, deliberate action, and taking it writes an entry against your account naming the columns you viewed. That log is what answers "who looked at whose data" when it is finally asked.
What changed since the last scan
Differences against the previous run, with a review workflow. Personal data appearing in a catalog that did not have it a fortnight ago is the finding that matters most, and the one a point-in-time audit is structurally incapable of producing.
Coverage, stated plainly
Which databases scanned completely and which did not. Where coverage is partial the console says so above the numbers rather than presenting a partial count as a total, and an entry is only ever recorded as no-longer-found for a database that scanned completely — because absence on a partial scan proves nothing.
Finding it is half. Proving you protected it is the rest.
Discovery is the part most programmes skip and every later obligation depends on. Protecting what you found, and evidencing that you did, are what the Rules actually score — and they sit on the same console rather than in three other tools that will disagree.
Hardening, measured against a benchmark
Reasonable security safeguards is the obligation carrying the largest penalty, and it is a measurement rather than an assertion. Each database is assessed against CIS hardening benchmarks with a pass rate, the observed value set against the expected one, and remediation text on every finding.
An access review you can hand to an auditor
Every principal on the database, what it can do, whether it can log in, and whether it has a password at all — reviewed daily, with a manual trigger for when you need it current rather than accurate as of this morning.
The evidence exports
DPDP sits beside PCI-DSS, HIPAA, GDPR and SOX in the compliance workspace, with control-level status as met, partial or gap and an explanation against each. Evidence reports download as JSON, per database, and are available to any signed-in user rather than gated behind an administrator.
Nothing leaves the database that should not
The monitoring account is read-only by design and its exact grants are published, so your DBA can read what is being asked for before approving it. Databases on a private network are polled by an on-site collector that reports outward, so no inbound route into your network is opened. Every query run through the console is recorded with the username against it, including the ones that were refused.
The difference when someone finally asks
A questionnaire and a spreadsheet
- The inventory is what application owners remembered in March.
- Aadhaar in a column called cust_ref was never in scope, because nobody looked at the values.
- "Who can read this table" is answered by asking the DBA.
- Access to personal data left no record, so the honest answer is that nobody knows.
- At breach time the first hours go on working out what was in the table.
PrahiX — discovery on a schedule
- The inventory is what the last scan actually found, twice a month.
- Findings are checksum-verified, so cust_ref is flagged on its contents.
- Entitlements are enumerated per account, including access inherited through a role.
- Every inspection is logged against a named user, with the columns viewed.
- At breach time the scope is already written down.
How discovery becomes evidence.
Four stages on a schedule. Each produces the artefact the next obligation asks for, which is the reason for doing them in this order rather than starting with the policy document.
Register the estate, read-only
Each database is added once with a dedicated read-only account — never an application login, and never root, sa or postgres. The grants required are published per engine, so your DBA approves a known list rather than a request to trust us. Publicly reachable servers are polled directly; anything on a private network is reached through an on-site collector, so no inbound route is opened. The connection is tested before it is saved.
Discover and classify
Discovery runs across the registered estate twice a month, with a manual trigger whenever you need it sooner — after a migration, before an audit, or during an incident. Columns are classified into DPDP categories from their contents, and where the identifier carries a check digit the number is validated rather than guessed at. A treemap sizes each table by how much sensitive data it holds, so the concentration is obvious before anyone reads the catalogue.
Attach the access picture
Entitlements, privileged sessions, failed logins and CIS hardening findings are gathered daily and land against the same databases as the personal data. That is what turns a list of columns into a risk statement: a table of Aadhaar numbers readable by a dormant account with no password is a different finding from the same table locked to one service principal.
Keep the record, export the evidence
Every inspection and every reveal is written to the access log as it happens, and scan-over-scan differences are retained with a review workflow. Control status and evidence reports export per database, so a question from a regulator, a customer or your own board is answered from what was already collected rather than from a fortnight of everybody's time.
What this replaces.
DPDP budget tends to go on describing the estate rather than on knowing it. That is backwards, and it is why the second audit finds what the first one missed.
| Capability | An annual consultant audit | Point tools plus a spreadsheet | PrahiX as the operating layer |
|---|---|---|---|
| Data inventory | Point-in-time, as at the audit date | Assembled by hand from schema exports | Rescanned twice monthly against live databases |
| How personal data is identified | Interviews and column names | Column names, mostly | Content classification, checksum-verified where the identifier allows |
| Coverage | Whatever was sampled | Whatever somebody connected | Stated per database, with partial scans declared as partial |
| Who can read it | Asked, and written down | A privilege query someone ran once | Entitlements enumerated per account, daily |
| Who looked at it | Not in scope | No record exists | Logged per inspection, with the columns viewed |
| Change since last time | Next year's report | A manual diff, if anyone remembers | Scan-over-scan differences with a review workflow |
| Evidence to hand over | A PDF describing the estate | Screenshots | JSON evidence reports and control status, per database |
| Cost shape | Per engagement, annually | Licences plus internal time | One operating subscription |
What PrahiX is, and is not
PrahiX finds personal data in the databases you register, tells you which accounts can read it, and keeps the record of who did. It does not make you DPDP compliant, and be wary of any product that says it does. Consent and notice, lawful basis, the consent-manager relationship, data-principal rights and grievance handling, retention and erasure decisions, and the accountability the Act places on the Data Fiduciary all stay with you. Scope worth knowing before a POC rather than during one: discovery covers databases you register, not files, mailboxes, object storage or SaaS; it runs on MySQL/MariaDB, PostgreSQL and SQL Server, while Oracle is monitored for metrics only and MongoDB produces none; and scans are scheduled twice monthly with a manual trigger rather than intercepting queries continuously. What changes is that the technical inventory every one of your obligations assumes exists, is current, and is evidenced.
- Checksum-verified discovery of Indian identifiers
- Column-level location, with entitlements
- Masked by default; every reveal is logged
- CIS hardening posture and JSON evidence
Trusted by operations teams across India

Have Questions? We've Got Answers.
The Digital Personal Data Protection Act, 2023 is India's general data protection law. It became workable when the DPDP Rules were notified in November 2025, and it commences in tranches across the eighteen months that followed. Consent-manager registration falls in November 2026. The obligations most organisations mean when they say DPDP — notice and consent, data-principal rights, reasonable security safeguards and breach notification — commence in May 2027. Published sources differ on the exact day within those months, so plan against the month rather than a date, and confirm the current position for your own obligations.
No, and be wary of anyone whose product claims otherwise. Compliance is an organisational obligation covering consent and notice, lawful basis, data-principal rights, grievance redressal, retention and erasure, and the accountability the Act puts on the Data Fiduciary. What a platform can do is tell you where personal data lives, which accounts can reach it and who has looked, and evidence that the technical safeguards were running. That is a large share of what an inquiry examines, and it is the share that cannot be produced retroactively.
A schema review reads column names. This reads sample values and classifies them, and where an identifier carries a check digit — Aadhaar, PAN, GSTIN, IFSC, card numbers — the number itself is validated rather than the label trusted. That is what catches Aadhaar stored in a column called cust_ref, which is the single most common way an inventory built from schema comes out wrong, and it is also what discards the part number sitting in a column called pan_no before it costs you a week.
No. Scans are read-only against your database, and a finding stores a masked sample rather than the value. The monitoring account is a dedicated read-only login whose exact grants are published per engine, so your DBA can review precisely what is being asked for before approving it. For databases on a private network an on-site collector polls locally and reports outward, so no inbound route into your network is required.
By default nobody. Inspect returns a masked sample, which is enough to confirm whether a finding is real. Unmasking is a separate action, and taking it writes an entry to the access log against your account naming the columns you viewed. Everything is scoped to your organisation — tenant separation is enforced on the server rather than in the interface — and every query run through the console is recorded with the username against it, including refused ones.
Twice a month, with a manual trigger for whenever you need it sooner — after a migration, before an audit, or in the middle of an incident. Entitlement review and CIS hardening assessment run daily. One habit worth forming: read the coverage banner before quoting a number, because where a database scanned only partially the findings are a floor rather than a total, and the console says so rather than letting you assume otherwise.
Discovery, security scanning and the SQL console run on MySQL and MariaDB, PostgreSQL and SQL Server. Oracle is monitored for metrics only — no discovery, no schema browser, no diagnostics. MongoDB can be registered and browsed but produces no metrics. We would rather you knew that before a proof of concept than found it during one.
No — it is one estate. The same monitoring, retention and evidence exports serve the RBI cyber security framework and the SEBI CSCRF cycle, with reporting scoped per entity and per period. DPDP adds the personal-data dimension on top of that same record, rather than a second set of tooling that will eventually disagree with the first in front of two different regulators.
Keep reading
Somebody is going to ask where the personal data is.
May 2027 is not far, and the inventory it assumes takes considerably longer to build than the notice does to write. Point us at one database — a replica is fine — and we will show you what a checksum-verified scan finds in it, and what the access record looks like from the day after.