Migration notes
We moved knock2.ai off Webflow. Here is why, and exactly how.
176 pages onto a Next.js static export we own, in the same repository as the rest of our product. The DNS switch took 35 minutes.
Two ways to read this
- Part one: why What staying put actually cost us. Start here if you are still deciding.
- Part two: how Ten steps with the specifics. Skip ahead if you have already decided.
One thing worth saying up front, because it changes what is realistic. Most of the extraction and porting was done by an AI agent working directly in the repository, with a human directing it and checking the output. Two years ago this was a contractor project measured in months. It is now a series of scripts you can read.
Part one
Why
Webflow was not the problem
Webflow is good. We would recommend it to plenty of companies.
The problem was that our website had become the slowest-moving thing we own.
We ship product daily. Marketing copy took a week. It lived in a visual editor that one person really knew, in a place our engineers could not see, behind a workflow nobody could review. Every test was a conversation. Every new page was a favor.
The money is the least interesting part
Start with the obvious line item, since everyone asks.
On Webflow's published pricing, a CMS site plan is $23 to $29 a month depending on billing, a Business plan is $39 to $49, and workspace seats are $19 to $49 each on top. For a site our size with a handful of people who need edit access, call it $70 to $150 a month, or $1,000 to $1,800 a year.
We would have kept paying that happily if it were buying us speed.
Now price the alternative most companies reach for instead. Contract development, depending on where you hire, runs anywhere from $10 to $150 an hour. Take the bottom of that range, with no agency markup on top.
Here is the thing. Even at $10 an hour, the money was never the constraint.
The constraint is the queue
A copy change means writing a brief, handing it off, waiting for someone's working day to overlap yours, reviewing what comes back, and asking for one more revision.
Best case that is a day. Realistically two. And it is two days whether the change is a new landing page or a single word in a headline.
So you start filtering. You stop asking for the small stuff, because it is not worth the coordination. You skip the test you are only 60 percent sure about.
That filtering is the real cost, and it never appears on an invoice. You cannot count the tests you did not run.
What changed for us is not that our developer got cheaper. It is that the queue disappeared.
| Before | Now | |
|---|---|---|
| Copy change | 1 to 2 days, via a handoff | about 2 minutes |
| New landing page | a ticket and a wait | same afternoon |
| Test nobody is sure about | usually skipped | just run it |
| Who can ship it | one person who knew the editor | anyone who can write a sentence |
| Reviewable, revertible | no | yes, like any other change |
The decision that de-risked everything
This is the one that matters most, and almost everyone gets it wrong.
You are rebuilding the site anyway, so the temptation is to redesign it at the same time. Resist that.
We split the two. The homepage, the customers page and all the CMS templates are new. Every other marketing page, including pricing, comparison pages and legal, was ported to render with Webflow's own stylesheet, so it looks pixel-identical to what was live the day before.
The reason is diagnostic. If your organic numbers move after a migration where you also redesigned 40 pages, you will never know which one caused it. Move the hosting while the pages look the same, let the numbers settle, then redesign with no migration risk attached.
Write down your current organic numbers before you touch anything. Whatever they are, they are your only way to tell a hosting problem from a design problem later.
The question everyone asks first: does this hurt SEO?
It is the right question, and the honest answer has four parts.
The platform is not the risk. The worry underneath this question is usually "React sites do not rank." That worry is really about client-side rendering, where a crawler asks for a page and receives an empty shell that only fills in once JavaScript runs. We do not do that. A static export writes one plain HTML file per URL at build time, which is the same shape of thing Webflow was serving. Ask our pricing page for its source right now and you get about 1,400 words of copy in the response, before a single script executes. Static export and server rendering are both fine. Client-only rendering is the thing to be careful about.
The real risk is the move, not the destination. Sites lose rankings in a replatform for a short and boring list of reasons. URLs changed. Redirects were dropped. Content quietly changed. Canonical tags went missing or pointed at the wrong host. A noindex left over from staging stayed on. The sitemap stopped matching what was actually published.
That list is finite, and a finite list is a checkable list. This is what each one maps to:
| What actually costs you rankings | What catches it |
|---|---|
| URLs change | Nothing was renamed. Parity first, redesign later |
| Redirects dropped | 31 rules in one file, host config generated from it, sync refuses a malformed rule |
| Content quietly changes | Every page compared against the old site on words, headings, images, title |
| Canonical missing or wrong | Build fails if any page lacks one pointing at the real domain |
| Staging noindex left on | One setting drives the header and all 176 tags; build fails if they disagree |
| Sitemap drifts from reality | Regenerated from what actually built, every time |
That is the real argument. Not "trust us," but "here is the list of ways this goes wrong, and here is the thing that fails the build when one of them happens."
Where owning it is genuinely better. Three things we could not do before.
- You can audit the whole site in one command. Every canonical, every title, every noindex, every internal link, across all 176 pages, in one pass. You cannot grep a hosted CMS. Two of the five problems in this article were found exactly that way.
- You can make the checks run themselves. A hosted CMS cannot refuse to publish because your sitemap disagrees with your robots tags. A build can, and ours does.
- You can actually ship experiments. The most underrated SEO advantage is the number of changes you get to make. Title tests, internal linking, a new page aimed at a real query: each of those used to cost a day or two of coordination, so most of them never happened.
And where it is worse, or at least not yet better. Two honest ones.
- Publishing without touching the repository is harder. We traded a visual editor for a pull request. That is a real cost, and it is the one thing you should weigh most carefully if nobody on your marketing team works in a codebase.
- We did not get faster on day one. The parity pages still load Webflow's generated stylesheet, because rendering under it is exactly what makes them identical. Page weight on those is roughly what it was. What changed is that the weight is now ours to remove, page by page, on our own schedule. On Webflow it was not a decision we could make at all.
What we did about it. Same URLs, because nothing was renamed. Every redirect carried. Content verified page by page against the old site. Canonicals asserted at build time. The sitemap regenerated from what actually shipped, then checked against the old one: of the 179 URLs Google already knew about, every single one either still builds at the same address or has a redirect pointing somewhere sensible. None were left without a destination.
That last check is the one we would do again first. It takes a few minutes, it is the difference between "we think the redirects are fine" and knowing, and it is the failure that costs you the most while being the easiest to miss.
Part two
How, in ten steps
The ten steps group into four phases. Nothing here is exotic. The order is what matters, because each phase is what makes the next one checkable.
01Mirror the live site
Before touching anything, take a complete copy of the site as it is actually served. Not as the CMS thinks it is. As a browser receives it.
We used wget:
wget --mirror --page-requisites --convert-links --adjust-extension \
--span-hosts --domains=yoursite.com,cdn.prod.website-files.com \
--restrict-file-names=windows --no-parent \
--wait=0.3 --random-wait
Roughly 5 to 15 minutes, and somewhere between 50 and 200MB. Ours came to 912 files.
This captures things the Webflow API cannot give you:
- The rendered HTML of every published page, which no API exports
- Inline embed elements and page-level custom code, as actually served
- Which scripts each page loads, including analytics and chat widgets you may have forgotten about
- CSS, JavaScript, fonts and images as served, including Webflow's generated stylesheet
sitemap.xmlandrobots.txt
Commit that mirror to your repository. It is your reference copy for everything that follows, and it is what you compare against when you want to know whether the new site actually matches.
02Pull the CMS out through the API
The mirror gives you rendered pages. It does not give you structured content, and you want both. A blog post in the mirror is HTML. A blog post in the CMS has a title, a slug, an author, a category, a publish date and a body, as separate fields.
We pulled 14 collections out of Webflow's API: blog posts, blog categories, customers, podcasts, integrations, integration categories, partners, partner categories, team members, careers, FAQs, plans, plan categories and SKUs.
The technique that made this practical: never let item bodies pass through the agent's context window. One blog post's rich text is around 10KB. We had 64 posts and 198 FAQ entries. Streaming those inline would exhaust the session before you finished.
Instead, the API responses get written to disk, and the agent parses them from disk with a script. The content never enters the conversation. This is the single most useful thing we learned, and it generalizes to any bulk data work with an agent.
Carry two extra fields on every item: the original CMS item id, and its last-updated timestamp. Those are how you detect drift. If someone edits a blog post while you are mid-migration, comparing timestamps tells you exactly which items to re-pull, rather than re-extracting everything and hoping.
Read your collection schemas first. Ours had traps. Blog posts had both a rich-text body and a plain-HTML body field, where the HTML one silently overrode the other when non-empty. FAQs had two different page-target fields that sometimes disagreed. Customers had a full story and a short card version, and you need both.
03Top up the assets the crawl missed
The page crawl fetches every asset a published page references. It does not fetch assets belonging to draft items, or to CMS entries no page currently links to.
Collect those URLs into a manifest while you extract the CMS, then fetch them in one pass:
wget --input-file=missing-assets.txt \
--force-directories --no-host-directories \
--directory-prefix=mirror
About a minute. Do it now rather than discovering a missing image in three weeks.
04Turn pages into data
Now convert the mirrored HTML into something your new site can render. We used cheerio, a server-side HTML parser, and wrote each page out as JSON.
The important decision is how much fidelity to keep, and we changed our minds about it partway through.
Our first attempt produced clean, neutral markup: stripped classes, semantic HTML, our own styles. It looked like a different website, because it was. We threw it away.
What we do now is full fidelity. Webflow's classes, ids, data attributes and inline styles are all kept. Each page's own <style> blocks and body class are carried verbatim. The pages render under Webflow's generated stylesheet, which we vendored into the repository, so they look essentially as the live site did.
That feels like cheating. It is actually the whole strategy. You are moving hosting, not redesigning. A page that looks identical is a page that cannot have broken.
Scripts get separated by purpose. Functional scripts stay: the Webflow runtime, custom code powering calculators and widgets, form embeds, structured data blocks. Tracking scripts are excluded, because the preview site must not fire analytics and pollute your real visitor data with your own testing. They come back as a deliberate step at launch.
Hidden elements stay hidden. Closed accordions, inactive tabs, mobile menus. With the real CSS and scripts present, those behave correctly. Strip them and you break the component.
Each page also gets its background and text color measured from the styled mirror and stored alongside the content. Two of our customer stories had been unpublished in Webflow after we ran the crawl, so that measurement found nothing and stored blanks.
The pages then rendered white text on a white background. They passed every structural check we had: right text, right images, right layout, completely unreadable. Someone had to look at them.
05Bring the assets in-house
Your pages still reference images, fonts and scripts on your old host's CDN. Every one is a dependency on a vendor you are about to stop paying.
Inventory them. We found 216, and gave each a flat local filename with a short content hash appended, rather than mirroring the CDN's directory structure. Webflow serves some URLs whose filename contains encoded slashes, and replaying that on disk invites a different encoding bug on every operating system and web server it passes through.
Then rewrite the references. Do it at load time, when content is read off disk, before anything renders. That matters more than it sounds. Our first attempt rewrote the built output afterwards, which misses cases where a long URL gets split across two chunks in the page payload. Neither half matches, neither gets rewritten, and the browser reassembles the original CDN URL and fetches it anyway. Six of our scripts were split exactly that way, and every visible check said we had localized them.
Three days after launch, while looking at something else, we found the site was still fetching a JavaScript library from Webflow's CDN on every page load. 86 references. It powers the scroll animations. Those animations work by starting content invisible and then revealing it, which means 27 pages had content that only appeared because a file on the vendor we were about to cancel answered the request.
Cancel Webflow, that file stops, those pages go blank. Not broken looking. Blank. No error, no failed build, nothing in any log.
If you take one thing from this guide: after you think you are done, search your built site for your old host's domain. Every match is a dependency you did not know you had.
06Carry the redirects
Export your redirect rules. Ours was 31 of them, accumulated over years, each pointing an old URL at its replacement.
Keep them somewhere readable with a note on each recording where it came from, and generate your host's config from that file rather than hand-maintaining both. We sync ours automatically before every build, so the two cannot drift.
Have the sync refuse to run on a rule that will not behave: a destination that is itself the source of another rule, which creates a chain, a rule pointing at itself, or a duplicate source where the second silently never fires.
Get this wrong and every old URL 404s on cutover day, and whatever rankings those URLs carried go quietly.
07Inventory the tracking tags
Open your old host's custom code panel and write down every tag, with its account or container id, and a decision beside each: carry, drop, or ask someone.
Ours came to a dozen. The inventory turned up two nobody could account for, set up by a former growth hire for tools we no longer used. Those got dropped.
On the new site the tags live in one configuration file that anyone can edit without touching code. It prints how many are enabled into the build log, so a deploy with analytics dark is visible immediately rather than discovered by a confused marketer a month later.
We did, and three tags from the carry list had never made it across. Two were measurement, which is annoying. The third was our chat widget, which meant that for about a day a visitor who wanted to talk to us had no way to do it.
That is not an analytics gap. That is a closed front door.
08Prove the new pages match the old ones
This is the step most migrations skip, and it is what makes the rest trustworthy.
We compare every built page against its mirrored original on measures that catch content loss: the h1, heading count, image count, word count, meta title and description. Word counts are compared as a ratio, so a healthy page sits near 1.00, and anything below 0.9 gets human eyes.
It runs after every build and writes a report. Ours reads 43 of 43 static pages and 132 of 132 CMS pages clean, and has stayed there through every change since.
You cannot eyeball 176 pages. You can read one number that tells you whether any of them lost content.
09Make the launch switch a single switch
Here is the one that would have quietly cost us the most.
While the new site was in preview, every page carried a noindex instruction so Google would not index something half-finished. Removing it was a line on our launch checklist.
Except it lived in two places: an HTTP header, and a meta tag on every page. The checklist named one.
Miss the second and the site goes live, every page loads, every form works, it looks perfect, and it ranks nothing until someone notices weeks later that organic traffic never came back.
So we replaced both with one setting, and made the build refuse to finish if the output disagreed with it. Flip one value and the header, all 176 meta tags and the check that verifies them move together.
We ended up with a handful of these:
- Every page must carry a canonical URL pointing at the real domain, because your preview URL stays publicly reachable and is otherwise indexable duplicate content
- The built site must reference nothing on the old host, asserted across all 216 assets
- The sitemap must not list a page that is noindexed, which is a contradiction Search Console reports as an error
If getting it wrong would still look right, a human will not catch it, so the build has to.
10The cutover
Ours was a weekend, deliberately, at the quietest hours we had.
The sequence for a subdomain like www is: add it as a custom domain on your new host, which gives you a DNS record to create. For us that record replaced the one pointing at Webflow, which means creating it is not preparation for the switch. It is the switch.
Then you wait for the certificate.
Everyone knows there is a gap between pointing DNS at the new host and the certificate being issued. What we did not know is that our old host had sent a browser instruction saying "only ever connect to this domain securely, for the next year."
So for returning visitors, that gap is not a warning they can click past. It is a hard block with no way through. First-time visitors got a dismissible warning. Anyone who had visited before could not reach the site at all.
Check whether your current host sends that header before you pick your cutover window. If it does, budget for a genuine outage, not a scare screen.
One more thing that will happen: your new host's dashboard will tell you almost nothing while you wait. Ours showed a red badge and offered exactly one action, "delete domain," with no way to distinguish "still working" from "gave up." The read-only API behind that same dashboard exposed the certificate's actual state, when DNS was last checked, and the specific challenge it was waiting on. That turned a guess into a fact and stopped us deleting a certificate request that was 20 minutes from completing. Find the API before you need it.
Leave the old site published and paid for. Ours stays up for a few weeks. Rolling back is putting one DNS record back, which is only true while the old site still exists.
Afterwards
Would we do it again
Yes, with the same sequencing.
The parts we found hard were not the parts we expected. Porting 176 pages was mechanical. Extracting the CMS was mechanical. The hard parts were the invisible dependencies, the things that fail silently, and the certificate window nobody writes about.
And the part worth internalizing if you are going to do this with an agent: generation is cheap now. Writing the extraction script, porting the pages, building the components, all of that is fast. Which means your judgment has moved almost entirely to verification. Every hour we spent on checks that fail the build paid for itself. Every problem we found afterwards was one where the only possible check was a person looking at the page.
Decide which of your failures are silent. Automate those. Look at the rest with your own eyes.
Knock2 identifies the people visiting your website, scores them against your ICP, and routes them to your team while they are still in market. The site you are reading this on is the one described above.
