RecoverSaves

RecoverSaves / saved_posts.csv format

What's really in saved_posts.csv

Reddit's export is complete and almost unreadable. Two columns, no titles, no subreddits, no dates. Here is exactly what is in it, and what turning it back into something you can read actually involves.

Two columns, and that is all

Tested against real exports. The file has exactly two columns:

ColumnContents
idThe thing's base-36 id, for example 1vxdkxs. Seven characters at time of writing, historically shorter.
permalinkA bare URL, for example https://www.reddit.com/r/somesub/comments/1vxdkxs/some_slug/. No title field, no body, no timestamp, no subreddit column.

How to request the export, and what else is in the ZIP, is covered separately. Saved comments arrive in a separate saved_comments.csv with the same two-column shape. The important property of both: they come from Reddit's own records rather than from the saved page, so they include everything past the roughly 1,000 the interface will show you. That is the whole reason the export is worth requesting.

The slug is not the title. A permalink often contains a lowercased, underscored fragment of the original title, and it is tempting to treat that as the title. It is truncated, it loses punctuation and capitalisation, and for a comment permalink it belongs to the parent post, not to the thing you saved. It is a routing artifact, not data.

Posts and comments look almost identical

Both kinds appear as ordinary Reddit URLs, and the difference is structural rather than obvious. A post permalink ends at the post; a comment permalink carries an additional id after the post's. That trailing id is the comment, and it is what you actually saved.

Reddit's internal naming reflects it: a post's fullname is t3_ plus its id36, a comment's is t1_ plus its own. Getting this wrong is the common failure in home-made scripts — a comment permalink parsed as a post resolves to the post's title, so the row looks successfully recovered while pointing at the wrong thing entirely.

id,permalink
1vxdkxs,https://www.reddit.com/r/somesub/comments/1vxdkxs/a_slug_here/
p6xb8ik,https://www.reddit.com/r/somesub/comments/1vxdkxs/a_slug_here/p6xb8ik/

Those two rows share a post id. The first is the post. The second is a comment inside it, and it is a different saved item with different text.

The rows that cannot be resolved

Not every line in the file points at something recoverable, and a tool that pretends otherwise will quietly drop them. There are four distinct shapes, and they deserve different treatment:

ShapeWhat it is
Share linkA shortened or share-formatted link that does not itself contain a thing id. It needs a network round trip to resolve to a real permalink, so it cannot be classified from the file alone.
Non-Reddit URLA link to some other host entirely. These do turn up in real exports.
Reddit, but not a thingA valid Reddit URL that is not a post or comment — a subreddit, a user page, a wiki page. There is no title to recover because there is no thing.
UnparseableAn empty cell, a malformed URL, or something that is not a URL at all.

These are reported rather than dropped. A row you cannot resolve is still a row you saved, and knowing there are eleven of them is more useful than an archive that silently contains eleven fewer items than your export did.

The file is a real CSV, and real CSVs are awkward

It is worth knowing what a parser has to survive here, because the naive spreadsheet approach quietly loses rows.

Rows with the wrong column count, an empty permalink, or a duplicate are each reported with the line number they occurred on, so a systematic problem reads as one finding rather than as hundreds.

Why recovering titles takes minutes, not seconds

The export has ids; the titles live on Reddit. Turning one into the other means asking Reddit about each saved item, and the honest constraint is that this must be done gently.

Reddit's info endpoint accepts up to 100 things per call, so a 5,000-item archive is around 50 calls rather than 5,000. Those calls are spaced by several seconds with a little jitter, and back off further if Reddit signals it is unhappy. A 5,000-item archive is therefore a few minutes of patient work, not an instant load.

That pacing is deliberate. Hammering the endpoint on your own logged-in session is how an account attracts a rate limit, and a faster tool that gets you throttled is not actually faster.

Turn that file into something readable. Feed saved_posts.csv to the extension and it rebuilds titles, subreddits and dates through your own session, flags what was deleted or removed, and shows you the whole archive — including everything past the 1,000 the saved page hides. Counting your hidden saves is free.

See how it works