Research ethics
Public data can still deserve careful handling. Something being technically retrievable says nothing about whether retrieving it is reasonable.
Built for research into publicly available Reddit content — not private information or identity tracing.
What this tool is for
- Journalism examining public discourse, coordinated activity or how a story spread.
- Academic research into online communities, with the usual ethical review that implies.
- Moderation research: understanding activity patterns in a community you help run.
- Community history, such as reconstructing how a subreddit changed over time.
- Personal archival research into your own account's public history.
What it is not for
Harassment, stalking, doxxing, identity tracing and discriminatory profiling are prohibited. So is building a case against a private individual. These are not edge cases we tolerate quietly — they are outside what the software will help with, and several of them are things it deliberately cannot do.
Boundaries built into the product
Policies that rely on good intentions tend to erode. Where possible the limits here are structural rather than stated:
- No feature estimates a person's location, timezone, age, gender, politics, religion, health or personality.
- The comparison tool compares aggregate statistics and will not tell you whether two accounts belong to one person.
- Interaction counts are defined precisely and framed as reply frequency. They are not contacts, friends or a social circle.
- Nothing reaches beyond Reddit archives. No cross-site username matching, no writing-style comparison, no reverse image search.
- Posting-time analysis carries a standing note that it does not reveal where someone lives, because that is the inference people reach for first.
An account is a person
A username with eight years of archived comments looks like a dataset. It is a person who wrote things over eight years, mostly without imagining they would be read as a single corpus. Volume creates a false sense of authority: a thousand records feels like a complete picture and is not.
Aggregation changes the character of information. Individually public facts, assembled, can reveal something none of them revealed alone. That is worth pausing over even when every input was public.
Interpreting carefully
- A timeline shows when records were captured. It does not tell you why someone posted.
- Word frequency describes text. It does not describe beliefs.
- Subreddit participation shows where content was posted, not what someone is.
- Silence in an archive is not evidence of silence on Reddit.
- A quoted comment may be one version of something later edited, or a fragment of a conversation you cannot see.
Before you publish
Verify against original sources where they still exist. Consider whether naming an account is necessary to your point, or whether the pattern is the point. Consider what happens to that person if you are wrong — and that archives are incomplete enough that being wrong is a live possibility.