Ask an architect how the practice website is performing and the answer usually describes how it looks. The work is judged by taste, and taste is what the practice sells, so this is a reasonable place to start. Visitors are doing something more functional. They are deciding whether this is the kind of practice they want, and whether to make contact.
What the site is for
A practice site has two jobs. It shows the work, and it produces the conversation that leads to a commission. The first is a design problem, and studios are equipped for it. The second is measurable, and most practices have never looked at it that way. Ask which project page produced last year’s enquiries and the answer is usually a referral, a publication, or a guess.
That leaves a gap in the studio’s own knowledge. The practice knows which projects it wants more of. It does not know which presentation of those projects brings the client in.
Why the redesign that “worked” may not have
The common sequence is to change the site and watch enquiries for a month or two. What that comparison measures is the practice’s reputation at that moment. A project gets published, a competition is shortlisted, an award lands, a former client refers someone, a quiet August turns into a busy September. Every one of those influences arrives in the same inbox count as the new homepage.
The number that moves most after a redesign is often the least informative: total visits. Traffic arrives from the publication, not from the layout. What the studio wants to know is whether a visitor who reached a project page then went looking for the contact details, and that question survives the noise only if it is asked in a way that isolates it.
What a controlled comparison is
A controlled comparison shows the current version to one group of visitors and the new version to another, at the same time, with assignment at random. Each visitor stays with the version they first saw. The only difference between the two groups is the change that was made.
Three rules keep the reading honest:
- Move one thing per test. Change the hero image and the enquiry link together and the result belongs to neither.
- Decide the metric in advance, and make it something close to the outcome you want.
- Leave the rest of the site alone for the duration, including the small edits that always seem urgent mid-test.
For a practice, the metric is usually a micro-conversion rather than a sale: reaching the contact page, opening the enquiry form, sending it, or clicking through to a project from the index. That is countable, and it moves faster than commissions do.
What is worth testing on a practice site
- The order of projects on the index. What visitors see first frames the rest of the practice.
- Case-study depth: the short summary against the full essay. Some clients want evidence of process, others want to scan.
- The hero image for a project: the photograph against the render, or the wide shot against the detail.
- Where the enquiry route sits: in the header, at the end of a project page, or in the footer.
- The number of fields in the enquiry form. Each one removed is a trade between qualification and completion.
- The commercial or multi-family page, if the practice wants that work and rarely gets asked for it.
Small traffic changes the method
Most practices do not have the visitors a retail test assumes. A studio with eight hundred sessions a month cannot run an experiment that needs four thousand people per variant, and pretending otherwise produces a result that is really a fluctuation.
Two honest options exist at that size. The first is to run tests only where the traffic is dense, usually the project pages with the strongest search presence, and accept that an experiment may take a quarter to report. The second is to stop counting and start observing: where visitors scroll, which pages they visit in sequence, and where they leave. Observation does not prove a cause, and it produces better hypotheses than a studio meeting does.
It is also worth setting expectations before the first test. In a meta-analysis of 1,001 A/B tests, 33.5% produced a statistically significant positive result, with a mean lift of 15.9% among the winners (Analytics-Toolkit). Two tests in three return nothing conclusive, and a programme survives on the premise that most attempts teach something and one in three pays for the rest.
When the practice also sells something
A growing number of studios sell products alongside their services: furniture lines, lighting, prints, hardware specified on projects. That side is a commerce site, and its questions are different. The price of a product can be tested directly, and so can the shipping threshold that decides whether a small order ships at all. The variables that matter on a product page, which is a different discipline from presenting a building, are the same ones a retail brand tests.
The benchmark figures are unflattering and useful. Littledata’s benchmark of 2,800 Shopify stores puts the average conversion rate at 1.4%, with anything above 3.2% placing a store in the best fifth (Littledata). Baymard’s aggregate of fifty studies puts average cart abandonment at 70.22%, which is the shape of the problem a studio faces the moment it adds a checkout to a portfolio site (Baymard Institute). A practice that treats its shop as an extension of the portfolio rather than as a store usually finds that the design sensibility carried over well and the merchandising did not.
Where the tool decides the outcome
Software built for commerce carries the mechanics that a studio would otherwise have to build: keeping one visitor in the same version across sessions, applying a tested price consistently from the product page to the order, and reporting a difference with the confidence behind it. Elevate A/B Testing runs price, page, shipping, image and split URL experiments on Shopify stores, with plans from $49 a month and a 14-day free trial, which is a low entry point for a practice testing whether its shop deserves more attention (Elevate pricing page, read 8 October 2026).
The same mechanics apply to the portfolio side, with one caution. A tool can run the experiment, and it cannot tell a practice which project to lead with. That judgement stays with the studio, and the test only reports whether the judgement was right.
A first test, in five steps
- Pick one project page with real traffic and one question about it.
- State the hypothesis as a sentence with a number attached: “leading the index with the housing work will raise clicks through to project pages.”
- Choose the metric, the sample size and the decision rule before launch.
- Run it for whole weeks, and keep other changes off the page.
- Write the result down, including the ones that changed nothing. The register is what stops the same discussion returning next year.
The short version
- A practice site shows the work and produces enquiries. The second job is measurable.
- A before-and-after comparison after a redesign measures the practice’s press, not its layout.
- Show both versions at the same time to randomly assigned visitors, one change at a time.
- At practice-scale traffic, test only the busiest pages and expect a slow read.
- Expect roughly one clear winner in three tests, and treat an inconclusive one as a closed question.
- If the practice sells furniture, prints or lighting, the shop is a commerce site and deserves the same discipline.
Sources
- Analytics-Toolkit, “What Can Be Learned From 1,001 A/B Tests?”, https://blog.analytics-toolkit.com/2022/what-can-be-learned-from-1001-a-b-tests/ (read 8 October 2026)
- Littledata, “Average Ecommerce Conversion Rate” (benchmark of 2,800 Shopify stores), https://www.littledata.io/ecommerce-conversion-rate (read 8 October 2026)
- Baymard Institute, “50 Cart Abandonment Rate Statistics” (average across 50 studies), https://baymard.com/lists/cart-abandonment-rate (read 8 October 2026)
- Elevate pricing page (read 8 October 2026), https://www.elevateab.com/pricing