When Customer Data Becomes a Security Risk

August 10, 2026

A US study of Home Depot customer profiles raises a broader question for retailers: how much confidence can security teams place in the data underpinning automated decisions?

Retail has become an increasingly data-driven business. Purchase histories, customer accounts, mobile apps, device identifiers, location information and third-party data are routinely combined to build detailed profiles of individual consumers. These profiles support personalisation, advertising, customer analytics and, increasingly, automated decision-making.

The commercial logic is straightforward: the more a retailer knows about a customer, the more accurately it can tailor products, services and communications. But that logic rests on an assumption that receives far less scrutiny — that the underlying data are correct.

A recent study by researchers from Northeastern University, New York University and Columbia Law School suggests that this assumption may be less secure than retailers would like to believe. Using Home Depot as a case study, the researchers examined personal information reports obtained by ten customers in four US states and compared the data held or inferred about them with their actual circumstances.

The sample is small and explicitly non-representative. It cannot establish how widespread such inaccuracies are across Home Depot’s customer base, still less across the US retail sector. What it does provide is an unusual glimpse into the scale and complexity of modern customer profiling — and the potential consequences when inference begins to diverge from reality.

Detailed Profiles, Questionable Accuracy

The reports contained substantially more than conventional account information. Alongside contact details and purchase-related data were stored or inferred attributes relating to household size, children, education, occupation, credit card use and geographic location.

Some of those inferences were strikingly inaccurate.

One participant in his early twenties, who lived with a roommate and had no children, was placed in a six-person household with young children and classified as belonging to a group of middle-aged urban renters with families. Another was identified as a Western European woman living in a seven-person household, although he was in fact a man living in a four-person household.

Occupational data were equally revealing. Six profiles included professional classifications; four were entirely wrong and two only partly correct, according to the researchers. In another case, a customer report contained more than 20 telephone numbers, ten email addresses and twelve physical addresses that the participant said did not belong to him.

The distinction is important. Some customer data are supplied directly: a name, email address or delivery address. Other attributes are inferred from patterns and signals. Those inferences may be commercially useful, but they do not carry the same evidential weight as verified information.

The researchers describe problematic examples as “junk inferences”: conclusions generated from data that appear specific and credible but may be based on weak or incorrect associations.

When Data Become Assumptions

Modern customer profiling depends heavily on combining signals from different sources. Purchase histories, browsing behaviour, devices, household information, location data and externally sourced attributes can all be used to produce increasingly granular classifications.

The problem is not necessarily poor processing. A system may work exactly as designed and still produce an inaccurate result.

Identity resolution is one obvious pressure point. Retailers try to determine whether different online and offline activities belong to the same individual or household. A login, smartphone, IP address, store visit or app session may all become part of a single identity graph.

In practice, such relationships are rarely clean. Families share devices and Wi-Fi networks. People move home. IP addresses change. Phone numbers are recycled. Several people may use the same computer or delivery address. Once information from outside providers is added, attribution becomes more difficult still.

The result can be a profile that is highly detailed but built on uncertain associations. Worse, once such data are replicated across systems, repeated agreement may create the appearance of confidence even when several platforms are ultimately relying on the same flawed source.

More data, in other words, do not automatically mean better data.

The Security Implications for Retail

The study itself focuses on customer data collection and inference rather than fraud detection or cyber security controls. The implications for security teams, however, are significant.

Retail security increasingly relies on digital identity. Fraud systems, login monitoring, device intelligence and transaction-risk engines all use combinations of behavioural and contextual signals to determine whether an activity appears legitimate.

Such decisions are only as reliable as the data on which they are based.

If a device is associated with the wrong household, a suspicious login may appear familiar. If an inaccurate location or account history is treated as authoritative, legitimate activity may be flagged unnecessarily. Neither scenario is demonstrated by the Home Depot study, but both illustrate why data quality deserves a place in the security discussion.

For CISOs, fraud teams and identity specialists, the lesson is not that inferred data should be discarded. It is that organisations should understand the level of confidence attached to different types of information.

A customer-confirmed address is not the same as an inferred household characteristic. A verified device is not the same as one linked through probabilistic matching. Security architectures that fail to distinguish between the two risk treating assumptions as facts.

Retail Media Widens the Exposure

The issue becomes more material as retailers expand into retail media.

Home Depot operates its own retail media network, Orange Apron Media, part of a broader industry shift in which retailers monetise their first-party data through advertising, audience segmentation and campaign measurement.

The strategic attraction is clear. Retailers possess transaction data that many traditional advertising platforms do not: they know not simply what consumers browse, but what they buy.

That advantage also expands the data environment.

Customer information can now move between CRM systems, analytics platforms, marketing technology, retail media infrastructure and external partners. An inaccurate attribute may therefore travel well beyond the system in which it first appeared. It can be copied, enriched, scored and redistributed, while its original provenance becomes increasingly difficult to establish.

This begins to resemble a data supply chain. For security and governance teams, the relevant questions are familiar: where did the information come from, how trustworthy is the source, when was it collected, how was it transformed and where has it subsequently been used?

Retailers have spent years improving visibility into software dependencies. Comparable discipline around data provenance is becoming harder to avoid.

Correcting the Record

A further challenge is what happens when an error is found.

Correcting a value in a customer-facing account may solve only part of the problem. Copies can remain in analytics systems, data lakes, advertising platforms or partner environments. In complex architectures, data correction is therefore less a single administrative action than a propagation problem.

For security teams, some corrections may also deserve closer scrutiny.

An incorrect occupation or household classification is most likely a data-quality issue. A customer profile containing unknown telephone numbers, email addresses or devices may point to something else: poor matching, an integration error, account compromise or identity misuse.

That does not mean every discrepancy should become a security incident. It does suggest that organisations need clearer escalation routes between customer service, privacy, fraud and security teams when anomalies involve identity-related data.

The boundary between data quality and security is becoming less distinct.

Trust Becomes a Security Property

The Home Depot study does not prove that customer data across retail are broadly unreliable. Its sample is too small for that. Its value lies elsewhere: it demonstrates how difficult it has become to separate verified information from statistical probability and algorithmic assumption inside modern customer-data ecosystems.

That distinction matters because cyber security has traditionally focused on confidentiality, integrity and availability. In highly automated retail environments, a fourth concern is becoming increasingly important: trustworthiness.

It is not enough to know that a data set has not been stolen or maliciously altered. Organisations also need to know how much confidence they can place in the information it contains.

Retailers should therefore treat provenance, confidence levels and correction mechanisms as part of the wider security architecture. Verified customer data should not be weighted in the same way as inferred attributes. Probabilistic identity matches should remain probabilistic. External data sources should be traceable. And corrections should be capable of reaching the systems that rely on them.

The question for retailers is no longer simply whether customer data are secure.

It is whether they are secure enough to trust.

Related Articles

Drone defence starts with situational awareness

From power plants and airports to logistics hubs and industrial sites, Europe’s critical infrastructure is facing a new security dimension above the conventional perimeter. A visit by former German Federal Transport Minister Patrick Schnieder to LivEye’s Drone...

Share This