Pre-launch security audit

Exposed Files Checklist: What to Scan Before a Website Goes Live

Automated crawlers and vulnerability scanners index new domains within hours of launch. This checklist walks through what to check on your staging environment — before DNS points at it and before those scanners get there first.

Run this on staging, not just after launch. Once a site is live and indexed, exposures can be logged by scanners within hours. The pre-launch window is your best opportunity to close these gaps before they become records in a threat-intelligence database.

Why the staging window matters

Internet-wide scanners — including Shodan, Censys, commercial vulnerability databases, and automated bots — probe large blocks of IP space continuously and index new hostnames within hours of first seeing them. Once a URL returns a 200 response for /.env, that response can be logged, scraped, and retained even if you block the path the next day.

The pre-launch window — your staging environment before DNS cutover, and the first 24 hours post-launch — is the critical moment. Catching an exposed file before launch costs minutes to fix. Catching it after a scanner has already exfiltrated the contents means rotating all exposed credentials, reviewing access logs, and potentially notifying affected users. The gap between those two outcomes is entirely a function of when you check.

What to scan for

These are the ten categories of sensitive paths that ExposureCheck tests for. Each includes the most commonly exposed paths in that category:

1. Environment files

The highest-severity category. These files contain API keys, database credentials, and service tokens. They are frequently left in the web root by accident during deployment.

  • /.env
  • /.env.local
  • /.env.production
  • /.env.staging
  • /.env.example
  • /.env.bak

2. Version control directories

An exposed /.git/ directory gives an attacker read access to your complete source code history, including every file ever committed — even files deleted from later commits.

  • /.git/config
  • /.git/HEAD
  • /.svn/entries
  • /.hg/store

3. Database dumps and backups

Database exports dropped into a web-accessible directory are a common data breach vector. These files contain full table contents — user records, hashed passwords, payment references.

  • /backup.sql
  • /dump.sql
  • /database.sql
  • /db.sql.gz
  • /backup.tar.gz

4. Configuration file backups

Editors and deployment tools sometimes create backup copies of config files with discoverable suffixes. These retain the original file's credentials in plaintext.

  • /wp-config.php.bak
  • /config.php.old
  • /settings.py.bak
  • /application.yml.orig
  • /config.json.bak

5. Private key and credential files

Service account keys, OAuth credential files, and private key files placed in the web root give immediate unauthorized access to the services they authenticate.

  • /credentials.json
  • /serviceAccountKey.json
  • /private.pem
  • /id_rsa
  • /client_secret.json

6. Log files

Application and error logs in the web root expose stack traces, library versions, user activity, and sometimes API keys or session tokens written during debug logging.

  • /error.log
  • /debug.log
  • /app.log
  • /logs/access.log
  • /storage/logs/laravel.log

7. Source maps

JavaScript and CSS source maps reconstruct your original source from the compiled bundle. For apps with proprietary logic or sensitive comments in the source, serving maps publicly exposes that code.

  • /static/js/main.*.js.map
  • /assets/*.js.map
  • /dist/*.css.map

8. Directory listing enabled

If directory listing is enabled, any path without an index file returns a browseable file list. Check upload and media directories specifically.

  • /uploads/
  • /media/
  • /files/
  • /backup/

9. .DS_Store and IDE files

macOS .DS_Store files and editor project files reveal directory structure. Downloading /.DS_Store reveals the filenames in that directory even without directory listing enabled.

  • /.DS_Store
  • /.idea/workspace.xml
  • /.vscode/settings.json
  • /Thumbs.db

10. Admin paths and exposed APIs

Development admin panels, API documentation left accessible in production, and default CMS admin paths that haven't been relocated are found by scanners within hours of launch.

  • /phpmyadmin/
  • /adminer.php
  • /api/docs
  • /swagger-ui.html
  • /.well-known/

Reading HTTP status codes

When you probe a path manually with curl -I, the response status code tells you exactly what's happening:

StatusWhat it meansAction required
200 OK The file or path is publicly accessible. If the response body contains credentials or config, this is an active exposure. Fix before launch — block the path in server config or remove the file from the web root.
403 Forbidden The file or directory exists and the server is blocking access. The resource is present but protected. Better, but not ideal. A 403 confirms the resource exists to a scanner. Prefer a rewrite rule that returns 404, or remove the file entirely.
404 Not Found The path returns no content — either the file doesn't exist, or the server is hiding it via a rewrite rule. Correct result. Sensitive paths should return 404 — they reveal nothing about server structure.
301 / 302 Redirect, usually to a login page or custom error. Generally fine. Follow the redirect and verify the destination doesn't expose sensitive content.

Quick manual check with curl (use -I for HEAD requests — no body download, faster):

# Check specific paths on staging
curl -I https://staging.yoursite.com/.env
curl -I https://staging.yoursite.com/.git/config
curl -I https://staging.yoursite.com/wp-config.php.bak

# Check for directory listing
curl https://staging.yoursite.com/uploads/
# Look for "Index of /" in the response body

Blocking sensitive paths in server configuration

The right fix is a server-side block. Deleting files from the web root is better still — but if they might reappear on the next deploy (e.g., a .env copied by a deploy script), a server rule is the reliable safety net.

nginx

# Block dotfiles and common sensitive paths
location ~ /\. {
    deny all;
    return 404;
}
location ~* \.(sql|bak|backup|old|orig|log|gz|tar)$ {
    deny all;
    return 404;
}
location ~* /(credentials|serviceAccountKey|private|id_rsa)\.json$ {
    deny all;
    return 404;
}
# Disable directory listing
autoindex off;

Apache (.htaccess)

<FilesMatch "^(\.env|\.git|\.DS_Store|wp-config\.php\.bak|credentials\.json)$">
  Order allow,deny
  Deny from all
</FilesMatch>

<FilesMatch "\.(sql|bak|backup|old|orig|log|gz)$">
  Order allow,deny
  Deny from all
</FilesMatch>

# Disable directory listing
Options -Indexes

Vercel (vercel.json)

{
  "rewrites": [
    { "source": "/\\.env(.*)", "destination": "/404" },
    { "source": "/\\.git(.*)", "destination": "/404" },
    { "source": "/credentials.json", "destination": "/404" }
  ]
}

Netlify (_redirects)

/.env          /404.html  404
/.git/*        /404.html  404
/backup.sql    /404.html  404

Build artifact checklist

Modern JavaScript build pipelines (Vite, webpack, Next.js) introduce their own exposure risks if not configured carefully. Before launch, verify:

  • Source maps match your policy. In Vite: build.sourcemap = false (or 'hidden' for error tracking without public access). In Next.js: productionBrowserSourceMaps: false is the default.
  • .env files absent from build output. Run ls -la dist/ and ls -la build/ — no .env* files should appear there.
  • Environment variables embedded in the bundle. React/Vue/Svelte frameworks using VITE_ or REACT_APP_ prefixes embed those values in the JS bundle. Only public-facing values (analytics IDs, public endpoints) should use these prefixes — never private keys or secrets.
  • node_modules/ not deployed. Confirm node_modules/ is absent from your web root. This can happen with misconfigured rsync or FTP deployments that copy the entire project directory.
Inspect your dist/ directory before every deploy. Run ls -la dist/ and look for anything that shouldn't be public: .env files, *.sql dumps, or private key files. The same files in your project root can appear in build output if your deploy script copies the entire project folder instead of only the build output.

Running ExposureCheck on your staging site

ExposureCheck probes your URL for over 150 high-risk paths across all ten categories above — environment files, version control, database dumps, log files, source maps, admin paths, and more. It's the fastest way to cover the full list without manually curling each path individually.

Two moments to run it:

  1. Before DNS cutover — paste your staging URL (e.g., https://staging.client-project.com). Fix any 200-response findings before proceeding to launch.
  2. Within 24 hours of launch — run it again on the live domain to verify that production server configuration matches staging. Deploy processes sometimes differ between environments.

Scan my staging URL with ExposureCheck →

For secrets already checked into git history, also paste your source code into LeakCheck to find hardcoded API keys before they reach the server — ExposureCheck finds what's accessible at the URL level; LeakCheck finds what's embedded in the code itself.

Pre-launch workflow

  1. Run ExposureCheck on the staging URL. Paste the staging URL and review all findings. Any 200 response on a sensitive path is a launch blocker — fix it before proceeding.
  2. Manually curl the five highest-risk paths. /.env, /.git/config, /wp-config.php.bak (if WordPress), /credentials.json, and /backup.sql. Verify each returns 404, not 200 or 403.
  3. Check directory listing on upload and media directories. curl https://staging.yoursite.com/uploads/ — look for Index of / in the response. If present, disable with Options -Indexes (Apache) or autoindex off (nginx).
  4. Verify build output. Confirm source maps match your policy, .env files are absent from dist/, and no private key files were bundled with the deploy.
  5. Switch DNS and re-scan within 24 hours. Run ExposureCheck on the live domain to confirm the production server config matches staging. Production environments sometimes have different web server configurations.
  6. Check HTTP security headers. Exposed files and weak security headers are often found together. After confirming no exposures, run HardenCheck on the live URL to verify Content-Security-Policy, HSTS, and X-Frame-Options are correctly set.

Frequently asked questions

Can ExposureCheck scan a staging URL that's password-protected?

No. ExposureCheck makes real HTTP requests to your URL. If the staging environment is behind HTTP Basic Auth, a VPN, or IP allowlisting, the scanner cannot access it. Run the scan after temporarily removing the password restriction, or run it on the live domain after DNS propagation. The two-stage workflow (staging scan + post-launch scan) is the most reliable approach.

What's the security difference between a 403 and a 404 for /.env?

A 403 (Forbidden) means the file or directory exists and the server is blocking access — the resource is present. A 404 (Not Found) means the server returns no content, either because the file doesn't exist or a rewrite rule hides it. From a security standpoint, 404 is preferable: it doesn't confirm to a scanner that the resource exists. A 403 on /.env is better than a 200, but the best configuration is a rewrite rule returning 404 for all sensitive paths.

Should source maps be accessible in production?

It depends on your application. Source maps make production debugging much easier and are generally acceptable for non-sensitive applications. If your frontend code contains proprietary algorithms, business logic with competitive value, or internal comments that reveal system architecture, serving maps publicly exposes that. The safest approach for sensitive apps is to disable source map generation for production builds, or upload maps only to your error-tracking service rather than your web root.

My server returns 200 for /.env but the file is empty. Is that a problem?

Yes, for two reasons. First, an empty .env on staging may not be empty on production — the path being accessible is the server configuration gap, regardless of current content. When you deploy to production, a populated .env may land in the same accessible web root position. Second, automated scanners that find a 200 response on /.env log the URL and retry periodically. Fix the server configuration to return 404 for that path.

What if I'm on shared hosting and can't configure server-level rules?

On Apache-based shared hosting (most cPanel environments), you can add rules via .htaccess in your web root without server-level access. Create or edit public_html/.htaccess and add the deny blocks shown in the server configuration section above. On Netlify or Vercel, use _redirects or vercel.json rewrite rules to return 404 for sensitive paths before your first deploy.

How often should I run ExposureCheck after the site is live?

Run it immediately after launch, then again after any significant server configuration change, CMS or framework update, or deployment that alters your document root structure. A monthly scan is a reasonable baseline for sites that change infrequently. For actively developed sites, add ExposureCheck to your deployment checklist — it takes 30 seconds and catches misconfigured deploys before users or scanners do.

Also in the Copper Bay Labs ship-safety suite

Pre-launch file exposure is one dimension of a secure launch. These free tools cover the others: