Publishing a page and getting a page discovered are separate events, separated in practice by weeks. On a site carrying forty district names that no map agrees on and two hundred equipment pages, they can be separated by never.
The complaint arrives in roughly the same words every time. The pages are live. They look right. They have been live for three months. Nothing ranks, and when you check whether the search engine has even seen them, half of them return nothing at all. Somewhere between the content management system and the index, the site quietly stopped being read.
This is the least glamorous problem in search and one of the most common, and it lands harder on Houston businesses than on most. The reason is structural rather than technical. A market with no zoning and no single commercial center produces sites with enormous geographic sprawl — a page for Aldine, a page for Channelview, a page for Deer Park, a page for the corridor that runs between them, plus a capability page for every service and a specification page for every piece of equipment. Nobody planned that architecture. It accumulated, one reasonable decision at a time, and it now consumes more crawling than the business can justify.
A page that exists is not a page that is found
A search engine does not receive your sitemap and dutifully read every entry. It maintains a queue, allocates a finite amount of fetching to each host, and reorders that queue continuously based on what it has learned about the site. Publishing adds a candidate to a queue you do not control. That is the entire transaction.
Three distinct things have to happen before a page can rank, and the vocabulary matters because the failures look identical from the outside.
- Discovery. The engine learns the URL exists — from an internal link, a sitemap entry or a submission. A page with no internal link pointing at it and no sitemap entry may never be discovered at all.
- Crawling. The engine actually requests the URL and receives a response. Discovery without crawling is common on large sites and is where most of the weeks disappear.
- Indexing. The engine decides the crawled page is worth storing and serving. This is a judgment about value, and it is the step no submission mechanism anywhere can force.
Crawl budget, and what quietly spends it
Crawl budget is the informal name for the fetching a search engine is willing to spend on one host over a period. It is not published, not configurable and not a number you will ever see. It is real nonetheless, and on a site of a few thousand URLs it becomes the binding constraint on how fast anything new gets seen.
The budget is influenced by how fast the server responds, how often content genuinely changes, and how much value previous crawling produced. That last factor is the one businesses control and routinely damage. Every fetch spent on a page that turns out to be a near-duplicate of another page teaches the engine that fetching this host is low-yield, and the allocation shrinks accordingly.
The usual consumers
These rarely announce themselves, because each one individually looks harmless.
- Filter and sort parameters generating endless URL variants
- Near-duplicate district pages differing by one place name
- Paginated archives running to hundreds of pages
- Redirect chains three and four hops long
- Soft 404s returning a 200 status with an empty result
What restores allocation
None of these are quick, and all of them compound.
- A fast, consistent server response under crawl load
- Fewer URLs, each genuinely distinct
- Correct status codes, including honest 404s and 410s
- Internal links that reflect what the business actually sells
- Sitemaps that list only what should be indexed
The arithmetic is unforgiving on a sprawling site. If an engine is willing to fetch a few hundred URLs a day from your host and your site presents four thousand URLs of which three thousand are variations on a theme, a genuinely important new page is competing for attention against your own dead weight. Pruning is not tidiness. It is the mechanism by which the pages you care about get read sooner.
Which geographic pages are real, and which are wishful
Every metro produces service-area pages. Houston produces the worst version of the problem, because the place names here are not administrative units. There is no zoning, the city boundary is irregular to the point of comedy, and the names people use — Spring Branch, Alief, Greenspoint, the Energy Corridor, Clear Lake, Bayport, Kingwood — are a mixture of neighborhoods, unincorporated areas, master-planned developments, ZIP-code shorthand and industrial districts. They overlap. Two of them frequently describe the same address. Some are separate cities that residents call Houston anyway.
The result is a page list that could expand indefinitely, with no natural stopping point and no authority to tell you where the line is. Add equipment and capability pages for industrial work — every certification, every service class, every asset type — and a mid-sized contractor can reach several thousand URLs without ever writing anything a customer asked for.
| Test | What you check | Real page | Wishful page |
|---|---|---|---|
| Demand | Impressions in the last 90 days | Shown repeatedly for its own name | Never shown, or shown for the parent term |
| Service | Jobs actually completed there | Recurring work, crews assigned | One job in three years |
| Content | What changes if you swap the place name | Projects, permits, conditions, references | Nothing but the place name |
| Distinctness | Overlap with a neighboring page | Different buyers, different work | Same address could match either |
| Link | Internal links pointing at it | Linked from services and navigation | Reachable only from a sitemap |
A page passing four or five of those tests deserves crawl budget. A page passing one deserves to become a paragraph on a page that passes five. The rule that survives contact with reality: a geographic page is real when its content would be wrong if you moved it fifteen miles. If swapping the name is the only edit required, you have one page written twice.
Sitemaps are an instrument, not a formality
Most sites treat the sitemap as something a plugin generates and nobody reads. Handled deliberately it is the cheapest discovery lever available, because it is the one place where you state, in machine-readable form, exactly which URLs you consider worth the engine's time.
Structure carries information. A single flat file of four thousand URLs says nothing about priority. A set of section sitemaps — services, geography, equipment, editorial — indexed from a parent file tells you where discovery is failing the moment you compare submitted counts against found counts per section. The Semalt Indexing Hub parses submitted sitemaps recursively to three levels of nesting and accepts up to 1,000 sitemaps in a single job, which is what makes a segmented structure practical rather than an administrative burden.
Sitemaps can be submitted as an uploaded file or as a URL, which matters more than it sounds. The URL route keeps the engine reading whatever your system currently generates. The upload route lets you submit a deliberately curated list — the ninety URLs that actually matter this quarter — without changing what the site publishes. For a pruning project, the second is the faster instrument.
Submission, and what it does and does not buy you
Beyond sitemaps, URLs can be pushed directly. The tracker carries a daily budget of 1,000 URLs per account, bulk submission accepts up to 10,000 URLs in a single batch, and delivery runs through the IndexNow API to GoogleBot and BingBot. Those three numbers define the realistic shape of any indexing program.
The sitemap route
A durable statement of which URLs belong in the index, re-read on the engine's own schedule.
- Best for structure and for removals
- Segment by section so failures are locatable
- Upload a curated file when pruning
The submission route
A request for attention on a specific URL now, delivered through IndexNow and logged per URL.
- Best for new and materially changed pages
- Bounded by the daily budget, so spend it deliberately
- Produces evidence when a page is fetched and dropped
What submission genuinely buys is time. A page that would have been discovered in three weeks can be discovered in days, which on a site publishing project pages or capability updates is the difference between being visible for a bid cycle and missing it. It also buys evidence: a URL submitted, visited and then not indexed has told you something specific about the page that no amount of waiting would have revealed.
- Submit new and materially changed pages. A rewritten capability page, a new service area you genuinely opened, a project page for work that just completed.
- Do not resubmit unchanged pages on a schedule. It consumes daily budget, teaches the engine nothing and produces the appearance of activity.
- Submit removals as deliberately as additions. Pruned URLs should leave the sitemap immediately; the log will show whether the engine has registered the change.
Reading the status of a batch without fooling yourself
A submission job produces a per-URL submission log rather than a single verdict: bot visit with timestamp, status, and error detail where one occurred, alongside live counters for submitted, found and failed. Those three counters are usually read as a success rate. They are better read as three separate questions.
| Counter | What it confirms | What it does not confirm | Act when |
|---|---|---|---|
| Submitted | The URL reached the queue | Anything about the engine's response | Count is below what you sent |
| Found | A bot visited and got a response | That the page was indexed or kept | Found stays far below submitted |
| Failed | The request did not complete | Whether the cause is the page or the host | Failures cluster in one section |
| Timestamps | When attention actually arrived | How often it will return | Gaps stretch across whole sections |
Clustering is the signal worth hunting. Scattered failures across a few hundred URLs are usually transient. Failures concentrated in one section — every equipment page, or every page below a particular directory — describe a structural fault: a template returning the wrong status, a canonical pointing somewhere unhelpful, a directory nothing links to. Section-segmented sitemaps make that clustering visible immediately, which is the practical argument for building them that way. Full logs export as CSV or JSON up to 10,000 rows; the rendered PDF is capped at 250 rows and is a summary format, not an audit trail.
AutoSEO — running the discovery cycle automatically
For a site whose URL count grew faster than anyone intended to manage.
- Automatic keyword discovery and prioritization. Which, applied to a geographic page list, is what tells you which place names carry demand.
- On-site AI suggestions. Including where pages duplicate one another closely enough to be competing.
- Automated link building. From a partner network of more than 230,000 websites.
FullSEO — with people deciding what gets cut
For consolidation work, where the decision to delete a page is a business decision.
- Human review before on-site changes ship. Necessary when forty overlapping district pages are being merged into six.
- Manual keyword selection with automatic fallback. So a thin list does not stall the campaign.
- Specialists, developers and writers alongside the automation. Consolidation is writing work before it is technical work.
Common questions
My pages are in the sitemap and still not indexed. What now?
Check whether they were fetched at all. If the log shows no bot visit, the problem is discovery and internal linking, and submission will help. If it shows a visit with a clean status and the page still is not indexed, the problem is the page, and resubmitting it will change nothing. Those two cases look identical in a ranking report and require opposite responses.
How many service-area pages should a Houston contractor have?
Fewer than the place names available, which is the only honest general answer. Start from areas where you have completed recurring work and where the analytics show impressions for the name itself. In practice most contractors here find that six to twelve genuine area pages outperform forty thin ones, because the twelve get crawled regularly and the forty get sampled.
Does IndexNow work with Google?
Submission through the IndexNow API reaches GoogleBot and BingBot. What it delivers is notification that a URL exists or changed. It is a faster way to be discovered, not a way to be indexed, and the distinction is the whole subject of this article.
Should I delete old pages or redirect them?
Redirect when a genuinely equivalent destination exists — the merged area page, the surviving capability page. Return a 410 when nothing equivalent exists. Redirecting dozens of thin pages to a homepage is worse than deleting them: it keeps them in the crawl queue and teaches the engine that requests to your host end in nothing useful.
How long should a batch take to show results?
Bot visits often appear within days, and the timestamps in the log tell you exactly when. Indexing decisions take longer and are not on a schedule. For campaign-level movement, four to eight weeks remains the realistic window, and reading a batch after five days as a verdict on the campaign will mislead you.
A calculation worth doing before anything else
Take an industrial services company in the eastern part of the county. Forty-one district and corridor pages, one hundred and sixty equipment and capability pages, a project archive of three hundred entries, and paginated listings adding another eight hundred URLs. Call it 1,900 crawlable URLs. Everything is in the sitemap, because the plugin puts everything in the sitemap.
Now apply the tests from earlier. Nine district pages describe areas with recurring work and their own search demand; the other thirty-two are one page written thirty-two times. Sixty equipment pages carry specifications a buyer would actually search; the remaining hundred are catalog stubs. The paginated listings should not be indexed at all. The honest URL count is closer to 450.
At 1,000 URLs per day, the pruned site fits inside a single day of submission budget with room to spare, and the same finite crawling the engine was already spending now lands on pages that answer a question. Nothing was added. The improvement came entirely from deciding which pages were real — which, in a county with no zoning and no agreed boundaries, is a judgment nobody but the business can make.
What to do first, and what to expect
The sequence matters more than the tooling. Inventory the URLs the site actually exposes, which is almost always more than anyone believes. Apply the five tests to every geographic page and every equipment stub, and write down the verdict rather than debating it. Consolidate and remove, with 410s where nothing replaces the page. Rebuild the sitemap as segmented files that list only survivors. Then submit, section by section, and read the counters per section rather than as one total.
If you want the discovery half of this run properly — sitemap jobs, submission batches, per-URL logs and the counters that show where a section is failing — the Indexing Hub is available in the Semalt dashboard, and the underlying crawl and index tooling sits behind the same filters as the analytics. The consolidation half is editorial work and stays with you. How we sequence the two for regional clients is set out across our service pages, with worked examples on the blog.