Pre-launch security audit
Exposed Files Checklist: What to Scan Before a Website Goes Live
Automated crawlers and vulnerability scanners index new domains within hours of launch. This checklist walks through what to check on your staging environment — before DNS points at it and before those scanners get there first.
- Why timing matters
- What to check for
- HTTP status codes
- Server configuration
- Build artifacts
- Run ExposureCheck
- Launch workflow
- FAQ
Why the staging window matters
Internet-wide scanners — including Shodan, Censys, commercial vulnerability databases, and automated bots — probe large blocks of IP space continuously and index new hostnames within hours of first seeing them. Once a URL returns a 200 response for /.env, that response can be logged, scraped, and retained even if you block the path the next day.
The pre-launch window — your staging environment before DNS cutover, and the first 24 hours post-launch — is the critical moment. Catching an exposed file before launch costs minutes to fix. Catching it after a scanner has already exfiltrated the contents means rotating all exposed credentials, reviewing access logs, and potentially notifying affected users. The gap between those two outcomes is entirely a function of when you check.
What to scan for
These are the ten categories of sensitive paths that ExposureCheck tests for. Each includes the most commonly exposed paths in that category:
1. Environment files
The highest-severity category. These files contain API keys, database credentials, and service tokens. They are frequently left in the web root by accident during deployment.
/.env/.env.local/.env.production/.env.staging/.env.example/.env.bak
2. Version control directories
An exposed /.git/ directory gives an attacker read access to your complete source code history, including every file ever committed — even files deleted from later commits.
/.git/config/.git/HEAD/.svn/entries/.hg/store
3. Database dumps and backups
Database exports dropped into a web-accessible directory are a common data breach vector. These files contain full table contents — user records, hashed passwords, payment references.
/backup.sql/dump.sql/database.sql/db.sql.gz/backup.tar.gz
4. Configuration file backups
Editors and deployment tools sometimes create backup copies of config files with discoverable suffixes. These retain the original file's credentials in plaintext.
/wp-config.php.bak/config.php.old/settings.py.bak/application.yml.orig/config.json.bak
5. Private key and credential files
Service account keys, OAuth credential files, and private key files placed in the web root give immediate unauthorized access to the services they authenticate.
/credentials.json/serviceAccountKey.json/private.pem/id_rsa/client_secret.json
6. Log files
Application and error logs in the web root expose stack traces, library versions, user activity, and sometimes API keys or session tokens written during debug logging.
/error.log/debug.log/app.log/logs/access.log/storage/logs/laravel.log
7. Source maps
JavaScript and CSS source maps reconstruct your original source from the compiled bundle. For apps with proprietary logic or sensitive comments in the source, serving maps publicly exposes that code.
/static/js/main.*.js.map/assets/*.js.map/dist/*.css.map
8. Directory listing enabled
If directory listing is enabled, any path without an index file returns a browseable file list. Check upload and media directories specifically.
/uploads//media//files//backup/
9. .DS_Store and IDE files
macOS .DS_Store files and editor project files reveal directory structure. Downloading /.DS_Store reveals the filenames in that directory even without directory listing enabled.
/.DS_Store/.idea/workspace.xml/.vscode/settings.json/Thumbs.db
10. Admin paths and exposed APIs
Development admin panels, API documentation left accessible in production, and default CMS admin paths that haven't been relocated are found by scanners within hours of launch.
/phpmyadmin//adminer.php/api/docs/swagger-ui.html/.well-known/
Reading HTTP status codes
When you probe a path manually with curl -I, the response status code tells you exactly what's happening:
| Status | What it means | Action required |
|---|---|---|
| 200 OK | The file or path is publicly accessible. If the response body contains credentials or config, this is an active exposure. | Fix before launch — block the path in server config or remove the file from the web root. |
| 403 Forbidden | The file or directory exists and the server is blocking access. The resource is present but protected. | Better, but not ideal. A 403 confirms the resource exists to a scanner. Prefer a rewrite rule that returns 404, or remove the file entirely. |
| 404 Not Found | The path returns no content — either the file doesn't exist, or the server is hiding it via a rewrite rule. | Correct result. Sensitive paths should return 404 — they reveal nothing about server structure. |
| 301 / 302 | Redirect, usually to a login page or custom error. Generally fine. | Follow the redirect and verify the destination doesn't expose sensitive content. |
Quick manual check with curl (use -I for HEAD requests — no body download, faster):
# Check specific paths on staging
curl -I https://staging.yoursite.com/.env
curl -I https://staging.yoursite.com/.git/config
curl -I https://staging.yoursite.com/wp-config.php.bak
# Check for directory listing
curl https://staging.yoursite.com/uploads/
# Look for "Index of /" in the response body
Blocking sensitive paths in server configuration
The right fix is a server-side block. Deleting files from the web root is better still — but if they might reappear on the next deploy (e.g., a .env copied by a deploy script), a server rule is the reliable safety net.
nginx
# Block dotfiles and common sensitive paths
location ~ /\. {
deny all;
return 404;
}
location ~* \.(sql|bak|backup|old|orig|log|gz|tar)$ {
deny all;
return 404;
}
location ~* /(credentials|serviceAccountKey|private|id_rsa)\.json$ {
deny all;
return 404;
}
# Disable directory listing
autoindex off;
Apache (.htaccess)
<FilesMatch "^(\.env|\.git|\.DS_Store|wp-config\.php\.bak|credentials\.json)$">
Order allow,deny
Deny from all
</FilesMatch>
<FilesMatch "\.(sql|bak|backup|old|orig|log|gz)$">
Order allow,deny
Deny from all
</FilesMatch>
# Disable directory listing
Options -Indexes
Vercel (vercel.json)
{
"rewrites": [
{ "source": "/\\.env(.*)", "destination": "/404" },
{ "source": "/\\.git(.*)", "destination": "/404" },
{ "source": "/credentials.json", "destination": "/404" }
]
}
Netlify (_redirects)
/.env /404.html 404
/.git/* /404.html 404
/backup.sql /404.html 404
Build artifact checklist
Modern JavaScript build pipelines (Vite, webpack, Next.js) introduce their own exposure risks if not configured carefully. Before launch, verify:
- Source maps match your policy. In Vite:
build.sourcemap = false(or'hidden'for error tracking without public access). In Next.js:productionBrowserSourceMaps: falseis the default. .envfiles absent from build output. Runls -la dist/andls -la build/— no.env*files should appear there.- Environment variables embedded in the bundle. React/Vue/Svelte frameworks using
VITE_orREACT_APP_prefixes embed those values in the JS bundle. Only public-facing values (analytics IDs, public endpoints) should use these prefixes — never private keys or secrets. node_modules/not deployed. Confirmnode_modules/is absent from your web root. This can happen with misconfigured rsync or FTP deployments that copy the entire project directory.
dist/ directory before every deploy. Run ls -la dist/ and look for anything that shouldn't be public: .env files, *.sql dumps, or private key files. The same files in your project root can appear in build output if your deploy script copies the entire project folder instead of only the build output.
Running ExposureCheck on your staging site
ExposureCheck probes your URL for over 150 high-risk paths across all ten categories above — environment files, version control, database dumps, log files, source maps, admin paths, and more. It's the fastest way to cover the full list without manually curling each path individually.
Two moments to run it:
- Before DNS cutover — paste your staging URL (e.g.,
https://staging.client-project.com). Fix any 200-response findings before proceeding to launch. - Within 24 hours of launch — run it again on the live domain to verify that production server configuration matches staging. Deploy processes sometimes differ between environments.
Scan my staging URL with ExposureCheck →
For secrets already checked into git history, also paste your source code into LeakCheck to find hardcoded API keys before they reach the server — ExposureCheck finds what's accessible at the URL level; LeakCheck finds what's embedded in the code itself.
Pre-launch workflow
- Run ExposureCheck on the staging URL. Paste the staging URL and review all findings. Any 200 response on a sensitive path is a launch blocker — fix it before proceeding.
-
Manually curl the five highest-risk paths.
/.env,/.git/config,/wp-config.php.bak(if WordPress),/credentials.json, and/backup.sql. Verify each returns 404, not 200 or 403. -
Check directory listing on upload and media directories.
curl https://staging.yoursite.com/uploads/— look forIndex of /in the response. If present, disable withOptions -Indexes(Apache) orautoindex off(nginx). -
Verify build output.
Confirm source maps match your policy,
.envfiles are absent fromdist/, and no private key files were bundled with the deploy. - Switch DNS and re-scan within 24 hours. Run ExposureCheck on the live domain to confirm the production server config matches staging. Production environments sometimes have different web server configurations.
- Check HTTP security headers. Exposed files and weak security headers are often found together. After confirming no exposures, run HardenCheck on the live URL to verify Content-Security-Policy, HSTS, and X-Frame-Options are correctly set.
Frequently asked questions
Can ExposureCheck scan a staging URL that's password-protected?
No. ExposureCheck makes real HTTP requests to your URL. If the staging environment is behind HTTP Basic Auth, a VPN, or IP allowlisting, the scanner cannot access it. Run the scan after temporarily removing the password restriction, or run it on the live domain after DNS propagation. The two-stage workflow (staging scan + post-launch scan) is the most reliable approach.
What's the security difference between a 403 and a 404 for /.env?
A 403 (Forbidden) means the file or directory exists and the server is blocking access — the resource is present. A 404 (Not Found) means the server returns no content, either because the file doesn't exist or a rewrite rule hides it. From a security standpoint, 404 is preferable: it doesn't confirm to a scanner that the resource exists. A 403 on /.env is better than a 200, but the best configuration is a rewrite rule returning 404 for all sensitive paths.
Should source maps be accessible in production?
It depends on your application. Source maps make production debugging much easier and are generally acceptable for non-sensitive applications. If your frontend code contains proprietary algorithms, business logic with competitive value, or internal comments that reveal system architecture, serving maps publicly exposes that. The safest approach for sensitive apps is to disable source map generation for production builds, or upload maps only to your error-tracking service rather than your web root.
My server returns 200 for /.env but the file is empty. Is that a problem?
Yes, for two reasons. First, an empty .env on staging may not be empty on production — the path being accessible is the server configuration gap, regardless of current content. When you deploy to production, a populated .env may land in the same accessible web root position. Second, automated scanners that find a 200 response on /.env log the URL and retry periodically. Fix the server configuration to return 404 for that path.
What if I'm on shared hosting and can't configure server-level rules?
On Apache-based shared hosting (most cPanel environments), you can add rules via .htaccess in your web root without server-level access. Create or edit public_html/.htaccess and add the deny blocks shown in the server configuration section above. On Netlify or Vercel, use _redirects or vercel.json rewrite rules to return 404 for sensitive paths before your first deploy.
How often should I run ExposureCheck after the site is live?
Run it immediately after launch, then again after any significant server configuration change, CMS or framework update, or deployment that alters your document root structure. A monthly scan is a reasonable baseline for sites that change infrequently. For actively developed sites, add ExposureCheck to your deployment checklist — it takes 30 seconds and catches misconfigured deploys before users or scanners do.
Also in the Copper Bay Labs ship-safety suite
Pre-launch file exposure is one dimension of a secure launch. These free tools cover the others:
Paste code or a config file and see exposed API keys and secrets flagged before they reach the repo — finds what's in the code, not just what's accessible at the URL level.
Security headers HardenCheckScan your live site's HTTP security headers — CSP, HSTS, X-Frame-Options — and get a grade with remediation steps for each gap.
Dependency risk DepCheckPaste your package.json and see vulnerable, abandoned, typosquatted, and risky-license packages before they go into production.
Check your site for WCAG accessibility failures and privacy-risk patterns before they become legal liabilities.