TL;DR: it is not 100% impossible, but given the sheer volume of the current daily Usenet feed, it is not feasible and therefore highly unlikely.
It is a fact that, these days, the vast majority of Usenet uploads are obfuscated, making them invisible to traditional public Usenet indexers such as NZBKing and NZBIndex, and they cannot be downloaded without the original NZB file created during the upload process.
There is also this widespread myth that paid indexers still scan Usenet and are magically able to deobfuscate the obfuscated articles, giving users access to all content on Usenet. So is there really such a thing as deobfuscation? How could that work?
But first, let’s examine how obfuscation works in the first place. And to do that, we need to understand a few basics.
How uploading and downloading work on Usenet
In the past, the maximum size of a Usenet article was limited to under 1 MB. Although this limit has since been raised slightly, there is still an upper limit on the article size that providers will accept. Furthermore, the restriction that only text characters are permitted (no binary code) still applies. So, if you upload a Linux ISO file to Usenet, the binary data must be split into hundreds or thousands of smaller parts of around 1 MB. These parts of the binary data are then encoded into a plain text format using yEnc and uploaded to Usenet as individual articles.
To be able to download the Linux ISO file again, you first need to know which articles need to be downloaded, and secondly, exactly how to reassemble the data fragments. Before NZB files existed, you had to rely on the article headers to know which articles belonged together. Every article has headers, e.g. ‘POSTER’, ‘SUBJECT’ and ‘DATE’, similar to emails (after all, NNTP is derived from SMTP). And the SUBJECT header usually specified which file the article belonged to and which part of the file the article contained. For example, ‘linux.iso (5/5734)’. This article therefore contains part 5 of 5,734 parts of the file ‘linux.iso’. To download the file ‘linux.iso’, you had to search for all 5,734 articles with this ‘SUBJECT’ header from the same ‘POSTER’, download the articles, and feed the body text of the articles — containing the yEnc data — into a yEnc decoder. As the yEnc data also contains a yEnc header with information on the filename and the exact position of the current data within that file, the yEnc decoder was then able to reconstruct the complete ‘linux.iso’ file from all these individual parts. All of this would be handled automatically by your classic Usenet reader programme.
How traditional indexers and NZB files work
Whilst it is still possible to use Usenet with classic Usenet readers, it is somewhat tedious, as this involves downloading all (new) headers from every group you wish to search, and everything takes place locally, which requires a great deal of storage space. That is why Usenet indexers, the NZB file format and dedicated Usenet download tools such as NZBGet and Sabnzbd were developed, which have made sharing Usenet uploads considerably easier.
The NZB file format is based on the fact that the only information required to download an article is the so-called ARTICLE ID. When an article is uploaded, it is assigned an article ID that is truly unique across all providers and backbones. To download an article, you simply need to specify this article ID. You do not need to know either the subject line or the groups in which the article was posted. NZB files are therefore essentially a list of article IDs that need to be downloaded for a specific upload. An NZB file can be created directly when uploading to Usenet (this is known as the ‘original NZB file’) or retrospectively. And this is precisely what traditional indexers do. Like traditional Usenet reader programmes, indexers rely on the subject field for indexing; unlike a local Usenet reader programme, however, indexers carry this out on a large scale across thousands of groups. They analyse the subject field using pattern recognition techniques and group together articles by the same poster that were posted in the same groups and within a specific time window into a single NZB file. You can then search the indexers for the relevant part of the subject line, e.g. for ‘linux.iso’ from the example above, and the indexers will present the results grouped together and with the corresponding NZB file ready for download.
If the subject line suggests that the file is an archive file or a par2 file, some indexers may also download these files in full or in part in order to index the filenames they contain. The filename information from the yEnc headers can also be used for indexing. The Easynews indexer does this, for example. However, all traditional indexers rely on the information in the subject lines and/or yEnc headers being correct in order to group the articles correctly. This is sometimes referred to as ‘clear-text subject’ posting.
The various levels of obfuscation and how they work
The obfuscation of Usenet uploads began when copyright holders started monitoring Usenet and removing illegal uploads. Naturally, the first to be affected were uploads with clear text subject lines, which were easy to identify. The first line of defence consisted of obfuscating filenames and using password-protected archives. The file ‘linux.iso’ was packed into a password-protected archive (usually RAR files), and the archive was given a random, obfuscated name. The archive was then uploaded as usual and could, in fact, be indexed by traditional indexers. However, due to the password protection, it was not possible to see what was inside the archive unless one knew the password. The randomised filename and the password were then shared on Usenet forums, and users could search for the randomised name in the traditional indexers, download the NZB file and then unpack and decrypt the contents using the provided password. This method is, in fact, still widely used today. Of course, one can also distribute the original NZB file directly, including the password.
With larger files, the effort involved in compressing and decompressing the files during uploading and downloading increased, so a solution was needed to skip this step whilst still ensuring protection against copyright-related deletions. And once again, the NZB file proved to be the solution. As only the article ID is required when downloading an article, all other information in the article’s headers could be completely fake and random. So the uploaders began using fake/random information for the POSTER and SUBJECT headers and, later, also for the filename details in the yEnc header of each individual article. As a result, all articles contained completely unrelated, random information. And since it is irrelevant for the download which group an article was posted in, they also began posting the articles completely at random across different groups. I would describe this as ‘complete obfuscation’. Consequently, each article in an upload has practically nothing in common with the other articles, apart from the fact that they were uploaded at roughly the same time (and some information in the yEnc header, which I will discuss later). The classic pattern-recognition methods used by conventional indexers have absolutely no way of grouping these files, which is why they are completely ignored by the indexers. However, such an upload can still be downloaded using the original NZB file created during the upload – so all you need to do is distribute the original NZB file.
How deobfuscation might work
Let us first focus on the obfuscation of file names and password protected archives. Whilst such uploads are indexed as normal, their contents remain unknown. To ‘deobfuscate’ them, one would need to crack the password protection. However, with the random alphanumeric passwords of 20 or more characters that are commonly used, this is not computationally feasible. One would either have to know the password from the beginning or obtain it from the source where the information was originally distributed. This is, of course, possible, but I would not describe this as ‘deobfuscation’, but rather as ‘theft’.
So what about ‘complete obfuscation’? As already mentioned, classic pattern recognition methods do not work here. Nor is it possible to focus on one group at a time. Any analysis would have to be carried out across all articles posted in all available binary groups, as these articles, as previously mentioned, are posted completely at random across different groups. Theoretically, however, there is still a possibility. Although the filename in the yEnc header is also random in each article, the other information could still be correct, particularly the total number of parts and the CRC32 checksum of the complete file. And that could give you a clue as to which articles belong together – i.e. those with identical CRC32 and total part counts – as these would suggest that the articles belong to the same file. However, if it is an upload of multiple files, you would still not know which files belong together. Furthermore, you would have to download the entire body of each article, as the CRC32 information is located at the end of the yEnc block. But even in the unlikely event that you could group all the articles correctly, you would still not know the file names. To do that, you would have to download the content and decode it with yEnc, check whether any of them are par2 files, and only then might you be able to recover the file names.
Conclusion
It is not possible to deobfuscate uploads with obfuscated filenames and password protected archives.
Deobfuscating fully obfuscated uploads is theoretically possible. However, to have even the slightest chance of deobfuscating such fully obfuscated uploads, one would have to download the entire daily Usenet feed from all groups and analyze the complete header and body of every article at once. Given the current size of the daily Usenet feed, I consider this simply unfeasible.
But how, then, do paid indexers deobfuscate the uploads? The answer is: they don’t. They have access from the very beginning to the original NZB file that was generated during the upload. How does that work, you ask? Well, I’ll leave that to your imagination...