Using pdfx in the Terminal and in CI: From the First Command to a Reliable Script
Use pdfx on the command line to merge, compress and watermark, process many files at once and chain tools into one request with a pipeline. For scripts and CI, learn what to rely on: exit codes, the --json report, standard input and output, and why a conditional filter tool with no match still exits 0.

Before you put pdfx in a script or in continuous integration (CI), remember two things. Exit code 0 does not always mean a file was written: a conditional filter tool with no match also exits 0. And --json has a different shape for one input than for several, so a script that parses only the files array finds nothing when just one file matched. pdfx is an HTTP client: files are uploaded to the server for processing, there is no local engine, and there is no offline mode. The full command reference is in the CLI guide on the developer page.
Install it, then merge two files
You need Node 20.3 or newer, or Bun:
npm install -g @pdf123/cli
pdfx --version
If you would rather not install globally, put npx @pdf123/cli in front of the command. With two PDFs in the current directory:
pdfx merge a.pdf b.pdf -o merged.pdf
On success, standard output prints merged.pdf. That was Merge. Before switching tools, ask for the fields rather than guessing parameter names:
pdfx describe watermark
Add Watermark (watermark) - Add text or image watermarks to PDF files
Input: 1 file (.pdf)
Result: file
Fields:
--watermarkText <value> Watermark text [default: PDF123]
--fontSize <number> Font size [default: 30]
min 6, max 200
(An excerpt; the fields also include --rotation, --customColor and others.) A field is just a command-line option:
pdfx watermark in.pdf --watermarkText DRAFT -o marked.pdf
pdfx list lists all 95 tools, --category security lists one category only, and --query watermark searches by word.
Several files go into a directory, and existing files are not overwritten
Give a single-file tool several inputs and each file is written into the directory named by -o as soon as it finishes, not held until the end. If one file fails the others carry on, and the exit code at the end of the batch is 1.
pdfx compress a.pdf b.pdf -o small/
Standard output from one run looked like this, and the order can differ each time:
a.pdf -> small/a.pdf
b.pdf -> small/b.pdf
An existing directory works as it is; if the directory does not exist yet, -o needs a trailing /, otherwise small is treated as a file name. Without -o, results land in the current directory under the name the server gives them; an existing file is not overwritten and the new one becomes a-1.pdf, a-2.pdf. A specific file name (-o same.pdf) replaces existing content, as you told it to. By default two files are processed at a time; change that with --concurrency. If you press Ctrl-C midway through a batch, the files already written stay and the exit code is 130.
Open an encrypted file with --input-password PASSWORD. When a batch mixes locked and unlocked files, use --password-for FILE=PASSWORD, which you can repeat.
Call tools separately for intermediate files, otherwise use a pipeline
To merge, then watermark, then compress, when you do not need the intermediate results on disk, fold them into one request with pipeline; the intermediate files are not downloaded and uploaded again. If you want the file from each step, call the tools separately. A pipeline produces a single file.
pdfx pipeline a.pdf b.pdf --step merge --step compress -o merged-small.pdf
# When a step needs parameters, describe the steps in JSON
pdfx pipeline a.pdf b.pdf \
--steps '[{"tool":"merge"},{"tool":"watermark","params":{"watermarkText":"DRAFT"}},{"tool":"compress"}]' \
-o out.pdf
A pipeline has at most 8 steps; a 9th step gets HTTP 400: at most 8 pipeline steps allowed. A tool that needs a second file (for example overlay-pdfs, which overlays another PDF) cannot be a pipeline step. pdfx rejects it before uploading, with exit code 2 and the message Tool "overlay-pdfs" needs a second file and cannot run as a pipeline step.
Use standard input only when you want to attach the command to a pipe. A file argument of - reads from standard input, and -o - writes the result to standard output:
cat report.pdf | pdfx compress - -o - > report-small.pdf
In this mode standard output carries only the bytes of the file; messages and errors go to standard error, so redirecting is safe. We checked it with a PDF of about 1 KB: standard output held a PDF of 1040 bytes, the same size as the file written with -o, and standard error was empty. When you read from standard input and give no -o, the result is named stdin.pdf, written to the current directory with a note on standard error.
Scripts should match the reason, not the message
| Exit code | Meaning | Examples |
|---|---|---|
| 0 | Success, or a conditional filter tool with no match | A merge succeeded; the filter-page-count condition was false |
| 1 | The request was sent but failed; in a batch, at least one file failed | Wrong password, not a PDF, timeout, nothing to return |
| 2 | Usage error, nothing was uploaded | Misspelled tool name, a pipeline step that needs a second file, a result type that contradicts the output file's extension |
| 130 | You interrupted it | Ctrl-C during a batch |
On failure, besides one line of message, standard error carries three lines: reason:, code: and hint:. An encrypted file with the wrong password produced this locally:
pdfx: HTTP 400: The password is incorrect.
reason: wrong_password
code: bad_request
hint: Check the password and try again.
The way you pass the password changes the first line of the message: unlock --password gives the line above. With --input-password on another tool, the server treats unlocking as an internal step and the first line is HTTP 400: Pipeline step 0 (security/remove-password) failed: The password is incorrect., while the reason: line is wrong_password in both cases. With no password at all the message is This PDF is password-protected. Enter its password. and reason is password_required. A script should match the reason: line. code is coarser, and a common value is bad_request.
A misspelled tool name exits with 2, and the message suggests similar names:
pdfx: Unknown tool "compres". Did you mean: compress, decompress-pdf? Run `pdfx list` to see all tools.
A result type that contradicts the file name also exits with 2, and it happens before anything is written. Writing the ZIP that Split produces to x.pdf:
pdfx: The result is application/zip but "x.pdf" has a .pdf extension; name it *.zip or write into a directory
code: output_mismatch
A timeout also exits with 1, with code set to timeout. The unit of --timeout is milliseconds and the default is 300000, that is 5 minutes, so --timeout 60 means 60 milliseconds, not 60 seconds.
When the condition is false, the exit code is still 0
One class of tools answers a yes-or-no question: is the page count greater than N, does the file contain some text, is the file bigger than some size. When the answer is yes, they return the input file unchanged; when it is no, they return nothing, and pdfx prints a line no match and exits 0. These filter tools exist only in the SDK, the command line and MCP; the website has no page for them.
This differs from another kind of empty result. When PDF to CSV finds no table in the PDF, the server returns 204 and pdfx exits with 1 and reports no_content: the tool meant to produce content and did not, and the reason is in Empty Table Export (204): Your PDF Probably Has No Columns. For a filter tool, "no match" is the answer it was supposed to give.
To read it in a script, use --json. A match prints the file that was written; no match prints { "matched": false }:
# Match: the file is written, and its info is printed
$ pdfx filter-page-count three.pdf --pageCount 2 --comparator Greater --json -o big/
{ "path": "big/three.pdf", "contentType": "application/pdf", "bytes": 2594 }
# No match: no file is written
$ pdfx filter-page-count three.pdf --pageCount 5 --comparator Greater --json
{ "matched": false }
To keep only the files with more than 2 pages, you can write it like this. We ran it locally on a.pdf (1 page) and three.pdf (3 pages), and only the latter was kept:
mkdir -p big
for f in *.pdf; do
if pdfx filter-page-count "$f" --pageCount 2 --comparator Greater --json -o big/ \
| jq -e '.matched == false' >/dev/null; then
echo "$f: skipped"
fi
done
jq -e '.matched == false' exits 0 when there is no match and 1 when there is. On a match the file has already been written into big/ by pdfx; the if only decides whether to print "skipped", not whether to write.
With a single input, --json has no files array
With several inputs, standard output under --json is a full report. input is an absolute path. processed counts the files whose request succeeded, including the ones with no match. unmatched is how many of those had no match, and their entries are { "ok": true, "matched": false } with no path. When every file in a batch has no match, the exit code is still 0. When you pass --idempotency-key to a batch, the key sent for each file is <key>:<index>; the mechanism is in Idempotency-Key: Safe Retries for PDF Jobs.
{
"processed": 2,
"unmatched": 0,
"failed": 1,
"files": [
{ "input": "/work/a.pdf", "ok": true, "path": "out/a.pdf", "contentType": "application/pdf", "bytes": 1040 },
{ "input": "/work/broken.pdf", "ok": false, "error": "HTTP 400: The file is not a valid PDF or it is damaged.", "reason": "invalid_pdf" },
{ "input": "/work/b.pdf", "ok": true, "path": "out/b.pdf", "contentType": "application/pdf", "bytes": 1027 }
]
}
With a single input, the single-file path is used: --json prints { "path": ..., "contentType": ..., "bytes": ... } and there is no files array. On failure standard output is empty and everything is on standard error. When you expand files with a glob, the script has to handle both formats, whether one file matched or several.
Collect the rules above into one CI step
This script only compresses and uses no filter tool. A compression failure exits with 1, and the step fails with it. A filter tool with no match exits 0, so CI does not fail just because there was no match; whether no match counts as a problem is for you to decide by reading --json, as in the filter section above.
If any file is rejected, this step fails and lists the file names and reasons in the log. The logic is the same as before: read the exit code first, print standard error on failure, then use jq to pull the reason out of the batch report. Without a directory argument it exits at once, so that "$1"/*.pdf is never expanded to /*.pdf.
#!/usr/bin/env bash
# Usage: ./ci-step.sh docs
docs="${1:?usage: ./ci-step.sh <directory>}"
mkdir -p out
pdfx compress "$docs"/*.pdf -o out/ --json > report.json 2> errors.log
status=$?
if [ "$status" -ne 0 ]; then
cat errors.log >&2
jq -r '.files[]? | select(.ok | not) | "\(.input | split("/") | last)\t\(.reason)"' report.json >&2
fi
exit "$status"
We ran it locally with four kinds of input:
| Files in the directory | Exit code | Log |
|---|---|---|
| Two good PDFs | 0 | none |
| Two good PDFs plus a damaged one | 1 | pdfx: broken.pdf: HTTP 400: ..., then a line broken.pdf invalid_pdf |
| A single good PDF | 0 | none; report.json is in the single-file format |
| A single damaged PDF | 1 | reason: invalid_pdf, code: invalid_document and a hint line; report.json is empty |
The question mark at the end of .files[]? in jq keeps the single-file report (which has no files) from raising an error. A single damaged file has no batch report and the reason appears only in errors.log, so keep both lines of output.
Three more things to check before you put this in CI. pdfx needs the network: files are uploaded to PDFX_API_BASE, which defaults to https://pdf123.xyz. For documents that must stay inside your network, run a service yourself and point this variable at it; see What Self-Hosting Actually Buys You (and What It Costs). When you need an identity, put the key in PDFX_API_KEY and not in command-line arguments, where it would stay in the process list and the logs. The version currently published is 0.1.0; in CI use npx @pdf123/[email protected] ... to pin it, so the output format and exit codes do not change with new releases and an upgrade becomes a change you make on purpose.
The package's page on npm is @pdf123/cli.