Process Substitution in Bash
As you become more proficient in administering Linux servers, the complexity of your commands will keep increasing. We already saw earlier how bash parameter expansion allows you to efficiently modify strings on the fly, without needing intermediate steps. Process substitution in bash takes this philosophy even further, eliminating many intermediate steps for certain kinds of commands that would otherwise be overly long or require their own script. What would have taken multiple lines and a lot more time is now just a single line.
The reason for these tools is that Linux administration relies on quick and efficient command-line operations, particularly those that are performed regularly. So let’s see what process substitution in bash can do for us.
Process Substitution for Temporary Files
Let’s see the different ways in which Linux uses the concept of temporary files in order to speed up multi-step processes.
Sorting and Comparing Two Files
Let’s say you have two files and you want to compare them after sorting them alphabetically. Let’s say we have two files – file1 and file2, containing the following content:
File 1:
apple
banana
orange
grape
And File 2:
banana
apple
kiwi
grape
Like this:

Now these two files are different. But merely running “diff” on them doesn’t do justice to their similarities. After all, they both contain three of the same fruits, but they’re out of order. If we want to understand the differences between the files, we need to do the following:
- Sort the files
- Compare the contents
To do this regularly, we’ll have to run the following commands:
sort file1.txt > sorted_file1.tmp
sort file2.txt > sorted_file2.tmp
diff sorted_file1.tmp sorted_file2.tmp
rm sorted_file1.tmp sorted_file2.tmp
Here, we’ve created two temporary files, “sorted_file1.tmp” and “sorted_file2.tmp,” with the sorted contents, and then we’ve run “diff” on the sorted files. Moreover, after the command completes, we’ve had to delete these temporary files, otherwise, they’ll remain and clutter up the system. In all, we’ve had to run four separate commands.
Here’s the output:
2d1
< orange
---
> kiwi

But what if there was an easier way? What if we could find a way to use the temporary files on the fly and feed them directly into the “diff” command? Consider the following command instead:
diff <(sort file1.txt) <(sort file2.txt)
This command uses a special syntax:
<(command)
This syntax makes “command” return its contents in a temporary file that Linux creates in the background. You don’t need to worry about what it’s called. The file name automatically replaces what was there before. So when Linux sees:
<(sort file1.txt)
It runs the command “sort file1.txt” and creates a temporary file as a placeholder. It does the same with the other sort command, and then the “diff” command runs, and it sees the two temporary files with their sorted content and compares them to give the same output as shown here:

Think of the syntax as little functions that return the output of their commands in a throwaway file name. After the command runs, the temporary files are removed, and you don’t need to worry about them. They won’t use up any space beyond the time that the command executes. With one single line, we have replaced four lines. And it runs fast too!
Now let’s see another use for process substitution in bash.
Multiple Filtering Commands for the Same Process Output
Let’s say we have a long-running command that generates a certain output. Now you want to filter that output via “grep” for two distinct strings, and store the results of the two in a file. There are two ways to do this normally.
- Run the command twice and filter the output each time
- Store the command output in a temporary file, and filter the output twice, then delete the temporary file.
For example, let’s say we take the same two files we had last time. Now, we want to check the file for two strings – “apples” and “bananas”, and store the outputs separately. Like this:
cat file1.txt | grep "apple" > apples.txt
cat file1.txt | grep "banana" > bananas.txt

This is fine if the source command is inexpensive and fast to run, like cat. But what if the output is a script that takes a long time to run, or is massively intensive? For example, what if the process is a huge backup? You can’t run a backup command twice! At least not without being tremendously wasteful.
The other option is to store the entire output of the command in a temporary file and then run the grep commands twice. Like this:
cat file1.txt > temp_data.tmp
grep "apple" temp_data.tmp > apples.txt
grep "banana" temp_data.tmp > bananas.txt
rm temp_data.tmp

This is a better solution for a long-running process, but like the earlier situation, it requires you to create a temporary file that uses up disk space and which you need to remember to delete afterward. It’s not a bad solution, it’s just awkward. And if you need to send the command to someone else, or run it multiple times, the whole process can get annoying, short of dumping the entire command into a script and running it.
Instead, we can use process substitution like this:
cat file1.txt | tee >(grep "apple" > apples.txt) >(grep "banana" > bananas.txt)
This is a different syntax compared to the first example. Here, we use the following syntax:
>(commands)
This syntax takes the output of a command and sends it as input via a temporary file to another command. The “tee” command allows us to split the output of a program into various streams, along with dumping it onto the standard output. By default, “tee” can only write into actual files, which we would then have to delete after the filtering is done. But with process substitution, we can feed the temporary files into our grep commands and not worry about cleaning up afterward. Everything happens in a single, neat command.
Why Process Substitution is Faster than Regular Commands
In addition to being less of a hassle to maintain, process substitution has the advantage of being significantly faster than the alternatives. For one thing, creating real files as placeholders involves disk I/O. And the larger the output, the greater the disk read/write activity. Process substitution, however, uses named pipes or file descriptors instead of actual files. The whole thing happens in memory without any need for disk I/O.
Another benefit is that process substitution allows bash to leverage parallel execution. So in the first example, bash can sort the files and evaluate the difference between them almost instantly, while if you created temporary files, it would have no option but to do everything sequentially.
For small files, there’s no great difference between process substitution and regular commands from a performance point of view. But if the source commands are resource-intensive, or the files are large, then it’s very noticeable. Particularly for large processes like backups, it’s essential to use process substitution to conserve resources.
Conclusion
Process substitution is like a super tool for advanced Linux admins who want to streamline their workflow as much as possible and squeeze every bit of performance out of their commands. Once you learn their uses, you’ll automatically start seeing uses for them everywhere and will feel compelled to leverage their advantages!

I’m a NameHero team member, and an expert on WordPress and web hosting. I’ve been in this industry since 2008. I’ve also developed apps on Android and have written extensive tutorials on managing Linux servers. You can contact me on my website WP-Tweaks.com!

Leave a Reply