Research Ethics
Public Online Research Without Crossing the Line Into Doxxing
The line is not public versus private. It is whether your work assembles a person or describes a pattern.
Subdex · 2026-08-22 · 3 min read
The usual defence of online research is that the data is public. Anyone can read a Reddit comment; reading a thousand of them is the same activity at scale.
That defence is weaker than it sounds, and the reason is worth understanding — not because public research is wrong, but because the "it's public" argument does not draw the line in the right place.
Aggregation changes the thing
A single Reddit comment is a public statement. A thousand comments by one account, sorted, counted and charted, is something the person never published. Each input was public. The output was not.
This is not a novel observation. It is why phone directories were unremarkable and why a database joining phone numbers to addresses to movements was not. Aggregation produces information that none of the inputs contained.
Legal systems have converged on similar reasoning under names like the mosaic theory. The intuition travels: the ethical question is not whether each piece was public, but what the assembled picture reveals.
A more useful line
Rather than public versus private, ask whether the work describes a pattern or assembles a person.
Describing a pattern:
- How did discussion of a policy change across a community over two years?
- Do posts linking to a domain get different engagement than text posts?
- What vocabulary distinguishes two communities?
Assembling a person:
- Where does this account holder probably live?
- What is their likely employer, age, health status or politics?
- Which other accounts are probably the same person?
Both use public data. Both are technically feasible. Only the first survives the question "what happens to a specific person if I am wrong?"
Being wrong is likely
That question deserves weight, because archive research is error-prone in ways that are easy to forget.
Coverage is incomplete, so you are working from a sample of unknown shape. Timestamps show when things were captured, not why. Removal markers carry no reason. Accounts change hands, get shared, get compromised. Sarcasm does not survive aggregation — a comment counted as expressing a view may have been mocking it.
An analysis about a community absorbs these errors. An analysis about a person concentrates them onto someone who cannot correct the record.
Practical constraints that work
Ask what the smallest sufficient claim is. If the point is that a coordinated campaign existed, you need to show coordination — not necessarily to name the accounts.
Prefer aggregates in output even when you needed individuals in analysis. Reading individual comments to understand a pattern is fine. Publishing a table of usernames is a separate decision requiring its own justification.
Do not chase the identity question. The moment the work turns from "what did these accounts do" to "who are these accounts," it has changed category — regardless of technique.
Notice when a tool is doing inference for you. Software that reports "most active hours" is one interface change away from reporting an inferred timezone. Software that reports subreddit participation is one step from inferring characteristics. Where a tool refuses to take that step, that refusal is doing real work.
Consider the target's position. Someone named in published research about their Reddit activity generally cannot correct it, did not consent, and may face consequences disproportionate to anything they did. This is not an argument against ever naming anyone. It is an argument for the bar being high.
What this is not
None of this argues against studying public online behaviour. Understanding how communities form, how narratives spread, and how coordinated manipulation works requires exactly this kind of research, and it is often in the public interest.
The argument is narrower: "it's public" establishes that you may look. It does not establish what you may build, or what you should publish. Those are separate questions, and answering the first does not answer the others.
Related tools
Related reading
Archive coverage varies and records may be incomplete. Verify important findings against original sources where available.