Skip to main content
Use this guide to check X-Robots-Tag header on PDF responses before you dig into broader indexing issues. Google supports the X-Robots-Tag HTTP response header for non-HTML resources, including PDF files, and Yandex documents the same control. An unintended noindex value can make an otherwise public PDF ineligible for indexing.

Prerequisites

You need:
  • the exact public PDF URL;
  • a terminal with curl;
  • access to the server, proxy, storage, or CDN configuration if a header has to change.
Run the checks against the public URL that users and crawlers receive, not a local copy of the file.

Step 1: Follow redirects and print headers

Start with a HEAD request:
--location follows redirects. Read each response block and focus on the final one. A problematic result can look like this:
The line that matters is x-robots-tag: noindex. The PDF is downloadable, but the response tells supporting crawlers not to index it. To send a regular request and discard the body, use:
This form is useful when you want to see the headers returned during a normal request. Don’t test only the URL before a redirect; the rule that counts is the one on the final response.

Step 2: Check for scoped directives

The header can carry a general rule or a rule addressed to a specific crawler. Capture every X-Robots-Tag line instead of stopping at the first one.
Read the complete result. If noindex is not intentional, remove it from the component that emits the final PDF response. Check server rules that target .pdf, directory rules, storage metadata, proxy behavior, and CDN response-header policies. Then run the command again against the public URL.

Step 3: Inspect canonical information

Google supports canonical declarations for non-HTML files through an HTTP Link header. Print both the robots and link headers:
A self-referencing PDF canonical can look like this:
When an equivalent HTML page is the preferred representative of duplicate or very similar content, the header can name that page instead:
Use an absolute canonical URL, as Google recommends. Keep it consistent with your internal links and other canonical signals. A canonical is used when search engines choose among duplicates; it does not compel indexing, and it is not a way to map unrelated content.

Step 4: Validate the document

Header fixes don’t repair an unusable file. Verify these document properties:
  1. Public access: the crawler should receive the PDF without a password or login.
  2. Encryption: Google says encrypted PDFs cannot be indexed.
  3. Searchable text: select and search for text inside the document. Google can index PDFs that contain text and may apply OCR to text in images, but OCR is no substitute for a reliable text layer.
  4. Yandex limits: Yandex indexes PDFs up to 10 MB. Text PDFs are indexed in full; for image-only PDFs, its documented recognition covers the first three pages only.
  5. Discovery: link to the stable PDF URL from a relevant, crawlable page.
If the export contains only page images, regenerate it with searchable text or apply OCR and check the order of the resulting text.

Step 5: Confirm the correction

Repeat the public request:
Confirm that the final response no longer contains an unintended noindex, that any Link canonical names the intended absolute URL, and that the returned resource is the expected PDF. If the live response is still unclear, SpeedyIndex can offer some workflow context. Save the final URL and the relevant header lines with the audit record. That gives you a compact baseline for comparing the same public response after later server, storage, proxy, or CDN configuration changes. Once the headers look right, check whether the PDF URL is indexed in Google. Keep in mind that Google’s Indexing API is not a method for ordinary pages such as PDFs; it is no replacement for crawlable links, valid headers, readable content, and coherent canonical signals.

Sources