RecoverSaves / saved_posts.csv format
What's really in saved_posts.csv
Reddit's export is complete and almost unreadable. Two columns, no titles, no subreddits, no dates. Here is exactly what is in it, and what turning it back into something you can read actually involves.
Two columns, and that is all
Tested against real exports. The file has exactly two columns:
| Column | Contents |
|---|---|
id | The thing's base-36 id, for example 1vxdkxs. Seven characters at time of writing, historically shorter. |
permalink | A bare URL, for example https://www.reddit.com/r/somesub/comments/1vxdkxs/some_slug/. No title field, no body, no timestamp, no subreddit column. |
How to request the export, and what else is in the ZIP, is covered separately. Saved comments arrive in a separate saved_comments.csv with the same two-column shape. The important property of both: they come from Reddit's own records rather than from the saved page, so they include everything past the roughly 1,000 the interface will show you. That is the whole reason the export is worth requesting.
Posts and comments look almost identical
Both kinds appear as ordinary Reddit URLs, and the difference is structural rather than obvious. A post permalink ends at the post; a comment permalink carries an additional id after the post's. That trailing id is the comment, and it is what you actually saved.
Reddit's internal naming reflects it: a post's fullname is t3_ plus its id36, a comment's is t1_ plus its own. Getting this wrong is the common failure in home-made scripts — a comment permalink parsed as a post resolves to the post's title, so the row looks successfully recovered while pointing at the wrong thing entirely.
id,permalink 1vxdkxs,https://www.reddit.com/r/somesub/comments/1vxdkxs/a_slug_here/ p6xb8ik,https://www.reddit.com/r/somesub/comments/1vxdkxs/a_slug_here/p6xb8ik/
Those two rows share a post id. The first is the post. The second is a comment inside it, and it is a different saved item with different text.
The rows that cannot be resolved
Not every line in the file points at something recoverable, and a tool that pretends otherwise will quietly drop them. There are four distinct shapes, and they deserve different treatment:
| Shape | What it is |
|---|---|
| Share link | A shortened or share-formatted link that does not itself contain a thing id. It needs a network round trip to resolve to a real permalink, so it cannot be classified from the file alone. |
| Non-Reddit URL | A link to some other host entirely. These do turn up in real exports. |
| Reddit, but not a thing | A valid Reddit URL that is not a post or comment — a subreddit, a user page, a wiki page. There is no title to recover because there is no thing. |
| Unparseable | An empty cell, a malformed URL, or something that is not a URL at all. |
These are reported rather than dropped. A row you cannot resolve is still a row you saved, and knowing there are eleven of them is more useful than an archive that silently contains eleven fewer items than your export did.
The file is a real CSV, and real CSVs are awkward
It is worth knowing what a parser has to survive here, because the naive spreadsheet approach quietly loses rows.
- Quoted fields with embedded newlines. A single record can span several physical lines, so counting lines does not count records.
- Doubled quotes as the escape for a literal quote inside a quoted field.
- Mixed line endings. Both LF and CRLF appear in the wild.
- An unterminated quote near the end of a file, which will swallow the remainder if a parser is not defensive about it.
- Duplicate rows, which should be collapsed rather than fetched twice.
Rows with the wrong column count, an empty permalink, or a duplicate are each reported with the line number they occurred on, so a systematic problem reads as one finding rather than as hundreds.
Why recovering titles takes minutes, not seconds
The export has ids; the titles live on Reddit. Turning one into the other means asking Reddit about each saved item, and the honest constraint is that this must be done gently.
Reddit's info endpoint accepts up to 100 things per call, so a 5,000-item archive is around 50 calls rather than 5,000. Those calls are spaced by several seconds with a little jitter, and back off further if Reddit signals it is unhappy. A 5,000-item archive is therefore a few minutes of patient work, not an instant load.
That pacing is deliberate. Hammering the endpoint on your own logged-in session is how an account attracts a rate limit, and a faster tool that gets you throttled is not actually faster.
Turn that file into something readable. Feed saved_posts.csv to the extension and it rebuilds titles, subreddits and dates through your own session, flags what was deleted or removed, and shows you the whole archive — including everything past the 1,000 the saved page hides. Counting your hidden saves is free.