Skip to content

Reddit Research

A Practical Introduction to Researching Public Reddit Discussions

A workflow for getting from a vague question to a defensible finding, and the places it usually goes wrong.

Subdex · 2026-08-22 · 3 min read

Most Reddit research starts with a question too vague to answer and ends with a number too specific to defend. The work in between is mostly narrowing the question until the available data can actually address it.

Start by making the question answerable

"What do people think about X on Reddit" cannot be answered. Reddit is not a population, the archive is not Reddit, and "think" is not a measurable quantity.

Turn it into something the data can speak to:

  • In which communities was X discussed, and when?
  • Did discussion of X rise or fall between two periods within this community?
  • What vocabulary accompanied X in this subreddit?
  • What did this specific thread contain?

Each is narrower and each is checkable. The narrowing is not a compromise — it is the actual research step, and skipping it is why so many findings collapse under scrutiny.

Scope before you search

Archives do not support searching all of Reddit for a keyword. Keyword search must be paired with a community, an author or a thread.

This looks like a limitation and is mostly a discipline. A keyword scoped to a community is a question about that community. The same keyword unscoped would return whatever the index happened to surface, which is not a sample of anything.

So decide the scope first: which communities plausibly discussed this, and over what period.

Take a small sample first

Load a hundred records before loading a thousand. A small sample tells you whether the query is finding the right thing, whether the community is the one you meant, and whether the period has coverage.

Most flawed analyses would have been caught by reading twenty records early. Volume makes patterns look solid before anyone has checked whether the records are what they appear to be.

Read some of the actual text

Aggregate figures are seductive because they arrive already looking like conclusions. Term frequency will tell you a word is common. It cannot tell you the word appeared in a quotation, or a joke, or a negation.

Before treating any aggregate as a finding, read a sample of the records behind it. This is the step most often skipped and most often the one that would have prevented the error.

Check the shape of the coverage

Look at the timeline before interpreting anything. If records cluster in certain years, ask whether that reflects the community or the archive. A cliff-edge drop is worth investigating before it becomes a finding — collection gaps look exactly like communities going quiet.

Write down what you did

Which archive, what query, when, how many records, whether the load completed. It takes a minute and is the difference between a finding someone can check and a number in a document.

Archives change underneath you, so "run the same query" is not a reproduction instruction. What you had, and when, is.

Stop at what the data supports

The final step is refusing the extra sentence. "Discussion of X in this community roughly doubled between 2022 and 2023 within the records I loaded" is supportable. "This community became increasingly concerned about X" is a claim about mental states that no volume of records establishes.

That last sentence is always tempting because it is the one that sounds like a finding. It is also the one that turns research about content into a claim about people.

Related tools

Related reading

Archive coverage varies and records may be incomplete. Verify important findings against original sources where available.