Logs are supposed to help you debug the system.
They are not supposed to become a second customer database.
Yet that happens constantly.
An API request gets logged. A request body gets captured. An exception includes a user object. An authentication token appears in a trace. A support message gets written to an error event.
The logging system then retains the data for months.
Why logs are different
Engineers know where the database is.
They often do not know all the places application telemetry goes.
A modern stack can include:
- Application logs
- API gateway logs
- CloudWatch
- Sentry
- Datadog
- OpenTelemetry traces
- APM tools
- Error tracking
- Debug exports
Some logs are structured. Some are free text.
The first rule: stop logging the payload
A common pattern is:
logger.info({ request, response })
Convenient.
Terrible privacy architecture if request or response contains personal data.
Instead, define explicit fields.
request_id
user_id
status
latency
error_code
Then include personal data only when there is a clear operational reason.
Search for high-risk fields
Look for:
- Phone
- Address
- PAN
- Aadhaar-related data
- Authentication tokens
- Session identifiers
- KYC documents
- Payment details
- User-generated text
Then inspect free-text logs, because names and addresses may appear without predictable field names.
Logging and DPDP retention
The final DPDP Rules specify a minimum one-year retention requirement for certain personal data, traffic data, and processing logs used for the security purposes specified by Rule 8(3), unless another law requires longer retention.
That does not mean every application log should be kept for one year.
It means teams need to understand which logs fall within the rule and which retention policy applies.
The difference matters.
Build a log privacy policy
For each logging system, record:
| Log type | Personal data? | Purpose | Retention | Access |
|---|---|---|---|---|
| API access log | Possibly | security / operations | defined policy | security |
| Error trace | Possibly | debugging | limited | engineering |
| Audit log | Yes, where applicable | accountability | applicable requirement | privileged |
Now you can reason about it.
Redaction before ingestion
The strongest control is upstream.
Do not send unnecessary data to the logging system and then depend on a filter to remove it later.
Use:
- Structured logging
- Field allowlists
- Redaction middleware
- Tokenisation
- Secret detection
- Request-body filtering
The test engineers should run
Create a synthetic account with distinctive dummy data.
Then trigger:
- Signup
- Login failure
- Payment error
- Support interaction
- API validation failure
Search the observability stack.
If the dummy personal data appears in five systems, you have found five privacy surfaces.
Common mistakes
Logging whole request bodies.
Keeping debug mode enabled in production.
Assuming vendor log retention matches yours.
Forgetting traces and error platforms.
No purge mechanism.
Where Privra fits
Privra can include observability systems in data discovery so personal-data mapping is not limited to business databases.
The goal is simple:
Your logs should tell you what the system did, without becoming a copy of everything the user ever sent you.