Prompt
North Central Loan's mainframe was compromised by Liber8tion. All we've been able to gather so far is this file, analyze it and figure what they were able to collect.
Walkthrough
In this challenge, you will analyze a mainframe file: specifically, a NETDATA file, commonly called an XMI file. XMI files are used to transfer sequential or partitioned datasets between mainframe environments. They consist of fixed-length 80-byte records with EBCDIC-encoded headers.
Mainframe computers have been produced from the early 1950s, with IBM being the main remaining manufacturer. They are incredibly reliable and can handle massive amounts of data and thousands of simultaneous transactions. Mainframes are still used widely by large enterprises, financial institutions, and governments.
Guide
In order to find what type of file this is, start by opening up the file in a basic text editor. It should show base64-like data like the following:
Recall the file extensions can be easily changed. We cannot trust the file name with the extension it was downloaded as. Running the file command will use the data inside the file (not the file extension) to classify what the file contains. Running file on the XMI shows that it is just “ASCII Text”. This is interesting since the file extension does not match that of a text file:
It’s possible that the file is encoded in base64 since that is what appears. Decode the file from base64 using the command line: base64 -d ARCHIVE.NETDATA.XMI > ARCHIVE.NETDATA.decoded.XMI
Lets check the out of file again. Now, the file command now shows it is an IBM NETDATA file:
In order to find what file extension is hidden in the XMI file, we need a way to extract the dataset.
Looking up information on “extracting netdata xmi” leads us to this python library as one of the results:
It’s necessary to create a virtual environment (venv) to install the tool. Creating this environment will protect the rest of the system
python3 -m venv venv
source venv/bin/activate
pip install xmi-reader
Using the documentation, extract the decoded XMI file with extractxmi ARCHIVE.NETDATA.decoded.XMI .
After running the command, the output may already have the correct file extension, in which case you can jump to extracting the contents. If the output has a .bin extension, run file on the extracted file to find what type it is. Rename the file to use the correct extension. You will now be able to extract the contents.
Unzip the extracted file with unzip ARCHIVE.NETDATA.DECODED.zip .
Examine the contents you just extracted.
What is that? We can use the file command again to learn more.
The file command identifies the format incorrectly. If we put EMAILS file through a PGP parser, it produces garbage. There just happens to be coincidental magic bytes that file is latching onto.
In order to better see this, lets verify what the file command is reading. Try opening the file in a hex editor.
0x40 bytes indicate the EBCDIC space character. Refer to resources above about XML IBM Mainframe files and how they are structured. This will give some hints about how to decode the content of these files.
Since the information is encoded, we are unable to read the content of these files. Use dd with the following format to convert the EBCDIC encoded files to ascii:
dd if=input.ebcdic of=output.ascii conv=ascii
The USERS file will most likely help us to answer the last questions about passwords.
USERS-asc now shows the hashed passwords in RACF format.
racf is plainly indicated in the password hashes. The hashes have been partially redacted.Next, we need to find what operating system (OS) uses this hash format. We can find that information via IBM’s own documentation: https://www.ibm.com/docs/en/zos/2.5.0?topic=users-passwords-password-phrases.
The password hashes are acquired and their type is known so we can solve the final questions by using a password cracking tool. Use John the Ripper because it has native RACF support : john --format=RACF --wordlist=/usr/share/wordlists/rockyou.txt USERS.txt - change USERS.txt to USERS-asc or whatever your decoded file is called.
To answer Question 6, you’ll have to add in the best64 rules list to help John crack it: john --format=RACF --wordlist=/usr/share/wordlists/rockyou.txt --rules=best64 USERS.txt
john --format=RACF --show USERS.txt will show the cracked passwords.
⚠️ By default, standard RACF passwords are case-insensitive as they are automatically converted to uppercase.
Additional Ways to Solve
To convert the files to ascii from EBCDIC, you could also use a script. A script is a way to a faster and more elegant output. These both give the decoded ascii outputs of ALL of the files at once.
With Bash using the dd command:
# this bash script loops through all the files and converts them from EBCDIC
LRECL=80
FILES=(EMAILS Q1 Q2 Q3 Q4 USERS)
for f in "${FILES[@]}"; do
dd if="$f" conv=ascii status=none \
| fold -w "$LRECL" \
| sed 's/[[:space:]]*$//' \
> "${f}-asc" # appends -asc to the filename, you can replace with .txt
echo "$f -> ${f}-asc"
doneWith Python:
Phil Young AKA Soldier of Fortran helped create this challenge.
Useful resources to solve this challenge:
- Mainframe: https://www.ibm.com/think/topics/mainframe
- XMI: https://xmi.readthedocs.io/en/latest/netdata.html
- Soldier of Fortran’s github tool for extraction: https://github.com/mainframed/Mainframed
- IBM Password Hashes: https://www.ibm.com/docs/en/zos/2.5.0?topic=setup-allowing-mixed-case-passwords-password-option
- Use our Tutorial Video below
Tutorial Video
Watch our full Tutorial Video to learn more specifics about mainframes, creating virtual environments (venv) and see a walkthrough of how to solve this challenge:
