Most sites here do not have a crawling problem. They have a decision problem disguised as one: nobody settled which pages should exist, in which language, and the crawler is left to sort out an address list that was never designed.
Publishing feels like an ending. The page is approved, it goes live, somebody checks the URL loads, and the task is marked done. From a search engine's point of view almost nothing has happened. The address still has to be found, fetched and judged worth storing before it can appear for anyone.
Between those steps sits a bottleneck that produces no error message. Pages go unvisited for weeks, a section launched in March is still absent in June, and nothing shows as failed because nothing failed — the crawler never got that far down the list.
Published, discovered and indexed are separate conditions
The three get treated as one event and are not. A published page exists on your server; a discovered page is one the engine knows the address of; an indexed page is one it decided to keep. Each transition fails on its own terms, for reasons unrelated to the step before.
Discovery happens mostly through links and sitemaps. A page two clicks from your home page is usually found within days. A project page reachable only through a filtered listing, or from a page that is itself buried, waits months, and one with no inbound path may never be found.
Indexing is a separate judgement, and it is a judgement. The engine asks whether the address adds anything to what it holds already. A French page repeating an English one sentence for sentence, a project entry differing from forty others only in a client name, a posting for a role filled last year — each is fetched, evaluated and quietly declined, and nothing in your logs records the refusal.
What actually spends the crawler's attention
Every site is allotted a rough working rate: how many requests an engine makes in a period, based on how quickly the server answers, how often the content changes, and how much of what it already fetched proved worth storing. There is no published figure and no dial.
The mistake is imagining the allowance goes to your important pages. It goes to whatever the crawler finds, in whatever order. On a mid-sized site with a portfolio, two languages and a decade of history, most fetches land on addresses nobody would defend in a meeting.
| What absorbs the fetches | How it usually arises | What it costs you |
|---|---|---|
| Filter and sort combinations | A portfolio or shop listing where every facet makes a new address | Thousands of URLs holding a few dozen items |
| A mirrored language tree | The whole site duplicated because bilingual felt like the correct default | Half the allowance spent on pages nobody queries |
| Tracking parameters in circulation | Campaign and referral tags shared, then linked back | One page fetched repeatedly under different addresses |
| Expired postings still resolving | Roles filled, pages left up and still linked | Fetches on content that is wrong as well as useless |
Server response time matters more than most teams expect. A crawler finding pages that answer in two hundred milliseconds takes more per visit; one hitting four seconds backs off, and it backs off across the whole site, not just the slow section. One expensive template throttles discovery everywhere.
The mirrored French tree, and the addresses it creates
The instinct here is to translate everything. It comes from a decent place — the company works in both languages, so the site should too — but it is applied as policy rather than as a series of decisions. Every English page acquires a French counterpart, the address count doubles overnight, and nobody asks which counterparts anyone would search for.
The result is predictable once stated plainly. A studio's forty project write-ups become eighty; a software firm's documentation doubles. Half the new addresses describe work for clients in California and London in a language those clients do not read, competing with the English versions that were doing the job.
Pages that should not exist twice
Content aimed at a market that searches in one language only, duplicated because the policy said so.
- Case studies for foreign clients
- Documentation with English-only terminology
Pages that genuinely need both
The handful of pages a client and a candidate will both open, which usually exist in only one language.
- Careers, team and studio culture
- The page explaining what the company does
Both errors sit in the same building at once: hundreds of translated addresses nobody queries, and three or four pages two entirely different audiences need in two languages, existing in one. The first drains the allowance, the second loses the enquiry, and neither shows in a crawl report because both work as built.
The useful exercise is uncomfortable and quick. Take the address list and answer one question per page: who searches for this, and in which language? Three answers are legitimate — English only, French only, or genuinely both — and the third is rarer than the policy assumes. A page with no answer should not be translated; it should be reconsidered.
- Sold abroad, written in English. Portfolio and capability pages aimed at studios and publishers outside Quebec rarely need a French twin nobody requests.
- Hired here, written in French. Recruitment, culture and location pages are read locally in French, and this is where a missing version costs you.
- Read by everyone, needed in both. The company description, contact details and main service pages carry both audiences and deserve two written versions.
- Answered by nobody. If you cannot name who searches for a page, translation is not the question; whether to keep it is.
Portfolios and project pages that never stop expanding
Creative and technical firms accumulate work pages faster than anything else on the site. Every finished project earns an entry: a game shipped, a sequence delivered, a collection photographed, a machining programme completed for a client who may not want naming. Nothing is removed, because removing work feels like denying it happened.
Ten years in, the portfolio is the largest section and the least differentiated. Entries share a template, a vocabulary and often whole paragraphs of boilerplate. Doubled into French, that section alone can account for most of a site's addresses.
The arithmetic is not exotic. Forty projects, two languages and six ways of filtering the listing produce several hundred addresses describing forty pieces of work. The crawler treats each as a candidate, fetches a good number, finds them nearly identical, and forms a view of how often your site is worth returning to.
The fix is editorial before it is technical. A dozen project pages written deeply enough to stand alone out-earn ninety near-identical entries, and the rest belong in a listing rather than at their own addresses. Filter combinations should not generate crawlable URLs unless somebody decided that combination is a page in its own right.
Postings that outlive the job
Recruitment pages are the clearest case of content with a shelf life, and where studios hire in waves they pile up fast. A twenty-person team can post thirty roles in three years. Each is a page, each linked from an index, several shared elsewhere, and almost none come down.
An expired posting is worse than a useless page. It is fetched like any other, competes with your current openings for the same terms, and tells whoever finds it something untrue about the company. In French, where the recruitment audience lives, that reaches exactly the people you want.
| Situation | Common handling | Better handling |
|---|---|---|
| Role has been filled | Page left up, quietly unlinked | Redirect to the careers index, permanently |
| Role recurs each year | A new address every cycle | One stable page, updated in place |
| Posting was widely shared | Deleted outright, links broken | Redirect, so accumulated links survive |
The same logic covers event pages, announcements and landing pages built for a launch that ended two years ago. None of it is dramatic; all of it spends fetches your current work needs.
The sitemap does a job nothing else does
A sitemap is not a ranking device and never was. It is a list of addresses you are prepared to stand behind, handed over rather than left for the crawler to infer from your navigation. Where deeper sections are reachable only through filters or paging, it is the difference between being found this month and being found eventually.
Which makes an honest sitemap a statement of editorial policy; the sitemap tools only report what you hand them. Containing every address your system can generate, it merely moves the problem. Containing the pages that ought to earn traffic, in the languages they ought to exist in, it describes what the site is for.
Submission, parsing and the queue
What the module accepts, how deep it reads, how much it holds at once.
- File upload or address. Submit a sitemap as a file, or point the module at its URL and let it fetch the document.
- Recursive reading, three levels deep. Index files pointing at index files pointing at sitemaps are followed three levels down, which covers any structure a site actually needs.
- Up to 1,000 sitemaps in one job. Enough for a large multilingual estate submitted as one operation rather than in pieces.
- Two jobs running, twenty waiting. Two sitemap jobs process concurrently and up to twenty more wait in the queue, so a migration is planned rather than fired off at once.
Three levels of nesting sounds like an implementation detail until a migration. Large sites split sitemaps by language, then section, then date, and such an estate is read correctly only if the parser follows the chain down. Where it stops at the first file, everything below is invisible and the submission still looks successful.
Keep language trees in separate sitemaps. It costs nothing and turns the submission into a per-language count you can read: this many French addresses offered, this many found, this many still waiting. Blended into one file, that information is unavailable.
Daily limits, batches, IndexNow and reading the log
Submission runs against fixed ceilings, and knowing them turns an intention into a schedule. The tracker works to a budget of a thousand addresses per day per account, while a bulk submission takes up to ten thousand in one batch.
The batch
Up to ten thousand addresses handed over in one operation, then worked through.
- Prepared once, submitted once
- Ten days of budget at full size
The daily budget
A thousand addresses a day per account, which is what decides your schedule.
- Spend it on pages you would defend
- Order the batch so it starts with those
Submission goes out over the IndexNow integration, a notification protocol rather than a request for a favour: it tells participating crawlers, GoogleBot and BingBot among them, that an address has appeared or changed. That shortens the wait before a fetch and has no bearing on what follows.
Reading the status of a batch is where the discipline shows. The log records, per address, whether a bot arrived and when, the status returned and the detail of any failure, alongside live counts of what was submitted, found and failed. Those three counters answer three different questions.
- Submitted tells you what you asked for. A count of your own actions, saying nothing about the outcome — and the number people quote in status reports.
- Found tells you the crawler arrived. A timestamped visit is real evidence: the address was reachable and the content fetched.
- Failed tells you what to fix today. Server errors, redirect chains and addresses that no longer resolve, each with detail attached. Short and actionable.
- None of them says indexed. Whether a page was kept is a separate question, answered by whether it starts appearing at all.
A high failure count is usually good news, because failures are specific and fixable. The harder pattern is a batch where everything was submitted, everything visited, and nothing appears. That is not a technical fault but the engine declining pages, and the answer is editorial: fewer addresses, written for somebody in particular.
Common questions
Should we submit our French pages as well, or only English?
Submit the French pages you would defend individually. If the French tree mirrors the English one completely, submitting all of it spends the daily budget on addresses that will not be kept. Submit what was written for a French-speaking audience, and use the exercise to decide what the rest is for.
How long should we wait before deciding a page will not be indexed?
If the log shows a bot visit and several weeks pass with no appearance for any query, treat it as declined rather than delayed. Resubmitting the same content produces the same result. Change the page, merge it into a stronger one, or accept that it need not exist.
We have ten thousand portfolio URLs. Where do we start?
Not with submission. Find out first how many of those addresses are filter combinations rather than pages, how many are translations nobody requests, and how many describe work you would still show a client. That usually cuts the list by most of its length, and the remainder fits inside the daily budget.
Does IndexNow make indexing faster or more likely?
Faster, sometimes. More likely, no. It informs participating crawlers that something changed at a specific address, which brings the visit forward. The decision to keep the page is made after the visit and is unaffected by how the crawler heard about it.
Does deleting old pages hurt us?
Removing pages that earn nothing and repeat each other generally helps, provided anything with accumulated links is redirected rather than dropped. Deleting outright is what hurts: broken addresses waste fetches, and links pointing at the old page are discarded with it.
The calculation, and the list worth writing first
Work an example through. A studio has 9,400 addresses: 3,100 English pages, 3,100 French counterparts produced by policy, and roughly 3,200 filter and tag combinations off the portfolio listing. At a thousand a day, submitting the lot takes ten days of budget, and it fits inside one batch of ten thousand with room to spare.
Now run the editorial pass first. The filter combinations are not pages and come out. Of the French counterparts, four hundred serve a French-speaking audience and the rest exist because everything was built. That leaves about 3,500 addresses worth submitting — three and a half days of budget, and every entry has somebody behind it.
The days saved are the small part. The real gain is that the crawler now meets a site where every address it fetches was chosen by somebody, which is what shifts its view of how often to return. The indexing module gives you the ceilings, the queue and the per-address log; the list you feed it is yours to write.
If the editorial pass is the part you keep postponing
The decisions stay yours; the execution can sit elsewhere.
- AutoSEO at 149 USD per month, per domain. Terms found and prioritised automatically, links built without your involvement, on-site recommendations and the analytics. Near 200 CAD — treat the conversion as rough.
- FullSEO at 500 USD per month, per domain. Terms chosen by hand with automatic fallback, placement held to a domain-rating threshold, changes reviewed by a person before release, and specialists, developers and writers behind it.
- Stream, the assistant in My SEO. One running feed of replies, reports, new links and open to-dos, bound to the project's data, and it accepts URL lists in bulk.
Start with a list, not a batch. Export every address, sort by section, and mark each with the audience it serves and the language that audience searches in. The exercise takes an afternoon and usually removes a third of the site. What is left goes through a bulk submission in one pass, and the structural work that follows has a defensible order. More of this sits on our blog.
To see the ceilings against your own address count, connect a property to a dashboard with your own domain attached and run a first batch through the log. The number worth watching is not how many addresses you submitted, but how many you could name an audience for — and in a company living in two languages, that number is always smaller, and better, than the one in the sitemap.