Let's Talk Privacy Graphs
Introduction
Privacy teams spend a lot of time figuring out how organizations work.
We're there to help teams make good decisions about personal data and fulfill our commitments to the people who entrusted us with it. To do that, we often need our organizations to answer a few familiar questions:
- What applications and services do we have?
- What data do they handle?
- Why do we have the data?
- How is it shared, retained, and deleted?
- What protections are in place?
The answers help privacy teams understand how organizations handle personal data. If we don't understand how data flows within and across organizations, how can we fulfill our privacy commitments?
To govern personal data effectively, we need to talk about graphs.
A Little About Graphs
Graphs are pretty simple: they're just things and relationships between them. In privacy terms, this means applications and the data flows between them.
For example, a service may be able to read from and write to a database. That service produces application logs, and the database has backups and replicas. Together, these connected resources form an application like so:

Organizations are made up of many connected applications that share data with each other. You need applications like an order service, a customer service platform, marketing, analytics—all the good stuff.
Privacy teams already know how these applications interact at a high level based on privacy reviews and institutional knowledge. We build ourselves a graph-based mental model of the applications and how they fit into the broader organization.
However, these mental models suffer from a few problems: they often aren't shareable, they come out of date as the business changes, and they aren't always correct. How we think an application works based on a privacy review may differ from how the application works in practice.
How could we build a better privacy graph?
Building a Privacy Graph
An effective privacy graph requires observable facts and organizational context. Privacy reviews provide context in spades, but where do we get more evidence?
In my experience, the best sources of facts are environments like AWS, GCP, and Azure, where the applications are actually running. For example, we can connect to our AWS accounts, discover relevant resources (compute, data stores, backups, etc.), and map the relationships and data flows between them.
This allows us to answer questions about what applications we have, what data they handle, and how that data is shared, retained, and deleted. From here, privacy reviews add the organization context like the relevant business purpose and the privacy policies and commitments that apply.
We can also update the privacy graph continuously as environments change. This makes it less dependent on either engineering teams' or privacy teams' understanding of the business.
It can flag new applications as they're created, as well as changes to existing resources that may be privacy impacting. We can identify new data flows between applications and also new uses of data, e.g., integrating with Amazon Bedrock.
The privacy graph serves as an evidence-based foundation for privacy programs.
Driving Better Outcomes
By automating more of the what, the privacy graph allows teams to focus on the why.
The privacy graph provides a common understanding of what applications exist, what resources they contain, and their data flows. In my experience, collecting this information takes up a large part of a privacy review (both in time and effort).
Why ask engineers for information that we can observe directly?
The questions at the beginning of this post don't change. However, by changing how we answer them, we can reduce our reliance and impact on engineering teams, empower privacy teams, and help the business move faster with confidence.
A privacy graph puts us in a better position to answer the question behind all the others: are we handling people's data in the way we've committed to?
That’s why we’re building the Privacy Graph at Truspecta.