# Brival — Full Content > Senior-led software engineering practice that modernizes existing software systems: stabilization, migrations, and rewrites delivered while the business keeps shipping, with AI-leveraged delivery built on over a decade of engineering discipline. Brival — Poland-based, serving clients across the EU and US. Contact: sales@brival.co. ## Services ### Code Audit URL: https://brival.co/services/code-audit A fixed-price audit of your code, infrastructure, and delivery process, from $5K, set by complexity and platform count. A founder-level engineer reads the system, not a junior with a checklist; you get a written report, a prioritized fix sequence, and a readout call. What you get: - A risk-ranked map of what you own: Code, infrastructure, tooling, and the delivery process, assessed against the same SDLC framework we build to, plus interviews with the people who run it. - The sequence, not just the findings: An ordered fix list, what first, what next, what can wait, not a diagnosis you’re left to triage yourself. - A report that speaks your scoreboard: Delivery metrics for an engineering leader, dollars and time for an owner, written to survive being forwarded to an investor or deal partner. - A readout call: We walk you through it, answer the hard questions, and you decide what happens next, if anything. Pricing: Starts at $5K. Fixed, all-in, paid once. The exact number depends on project complexity and platform count. — Fixed because a diagnosis shouldn’t be a negotiation: the price is set before you commit, and it stands alone, never credited toward follow-on work, because a diagnosis discounted by the cure stops being a diagnosis. ### Consulting URL: https://brival.co/services/consulting A recurring senior advisory engagement from 8 hours a month: fractional-CTO cover, running the department, training the team, and making your engineers faster with AI. We usually start with an audit, so the advice rests on findings rather than guesses. What you get: - What to track and how to run tech: A working management cadence, the numbers that matter for your stage, reviewed on a regular loop that sets goals and plans the work. - Fractional-CTO cover or team training: One senior brain across whatever the month needs, vendor and architecture decisions, hiring input, or structured upskilling of your team. - AI leverage from people who ship with it: How to make your engineers faster with AI on the codebase you actually have: rollout, measurement, and the prerequisites the tools depend on. Not slideware. - A monthly report: What was discussed, decided, and moved, a record legible to whoever you answer to: a board, an operating partner, a CFO. Pricing: Senior hours, from 8 hours a month. Recurring, with a 2-month notice period. We usually start with an audit first (fixed price, from $5K). — Hours plus a monthly report, because that’s what advisory is: senior attention applied where the month needs it, with a record of what it produced. Below the 8-hour floor nothing compounds. You’d be buying reassurance, not change. ### Project Development URL: https://brival.co/services/project-development A fixed-price, fixed-scope build: we commit to price, scope, and timeline in the contract, build to our written standard, and hand over with a 3-month warranty. New builds and modernizations alike. And because senior engineers build it with AI on a codebase held to that standard, work once quoted in years now ships in months. If it runs over, that’s our cost. What you get: - Software built to spec, then actually handed over: You end up owning everything: accounts, infrastructure, code, docs, plus training. It’s done when you can run it without us, not when it ships. - Modernization while you keep shipping: Incremental by default: staged cutovers, reversible steps, releases continuing throughout. We don’t pitch the clean-slate rewrite; when one’s right, the audit says so first. - Test automation in the price: Every feature ships with its tests in the fixed number, so regressions drop and stay down. (Escaped-defect rate is what we’re instrumenting across engagements, an intention, not yet a proven claim.) - A 3-month warranty: Bugs are on us for three months after release, with scheduled check-ins. Our incentive to build it right survives the final invoice. Pricing: Fixed price off the estimate, warranty included. Payments: up to 25% upfront, monthly invoices, final 10–20% after handover. — You pay more than the same build would on time-and-materials; that premium is the price of the risk sitting with us. In exchange: a number and a date you can plan around, and our incentive aligned with finishing, not billing. ### Team Augmentation URL: https://brival.co/services/team-augmentation Methodology-trained engineers embedded in your team, from 40 hours a month with a 1-month notice. Dedicated expert time immediately, no per-request scoping, no work orders, working to a craft standard that isn’t negotiable. What you get: - An expert on day one: Dedicated time inside your team, your cadence, your tools, no per-request scoping, no work orders between you and the work. - The craft standard, non-negotiable: Our people write tests, document as they go, and review properly, because it’s who we send, not because your process demands it. The standard is written down, so it doesn’t depend on who shows up. - An optional department install, no upcharge: Using hours you already pay for: we walk your delivery process, wire up the instrumentation (crash tracking, test tracking, a dashboard), and review it on a regular cadence. - Your team, faster with AI: Your engineers work next to people who ship with AI on a disciplined codebase and get faster the same way, capability that stays after we leave. The tools reward a craft standard; without one they multiply the debt instead. Pricing: Roles à la carte (developer, QA, PM) from 40 hours a month, with a 1-month notice. Engagements start at $5K a month; a full team typically runs $10–15K. — Monthly availability rather than an hourly rate card, because what you’re buying is a team that knows your product, not interchangeable hours. The department install carries no upcharge. The 1-month notice keeps it honest: if the value isn’t obvious, leaving is cheap. ## Methodology ### 01 Strategy & Business Context (delivery-flow) Strategy sets the intended business outcome: what gets built, why, and the number that should move when it ships. Every other block serves it; monitoring closes the loop by checking whether it did. AI stance: AI hands everyone a plausible strategy for the average case; what changed is that validating yours, with working demos instead of mockups, got cheap. ### 02 Requirements (delivery-flow) Requirements make sure the business can say what it wants, and verify it got it, before the budget is spent, not after. AI stance: The business-to-dev translation layer is shrinking: requirements work shifts from translating to deciding and verifying, and letting AI drive the product decisions is the first domino. ### 03 Architecture (delivery-flow) Architecture limits the blast radius: the structure that decides whether the system can grow, and whether one change ripples through everything it touches. AI stance: The architect matters more than ever: a wrong call at the start lands you in a hole months later, and AI accelerates the digging in both directions. ### 04 Coding (delivery-flow) Coding turns decisions into a system that stays changeable, code consistent and understood, so the next change is as safe as the first. AI stance: AI-first, with the engineer’s understanding non-negotiable: code nobody holds a mental model of is legacy code on day one. ### 05 Quality Assurance (delivery-flow) Quality assurance makes releases boring: nothing reaches users untested, and a bug fixed once stays fixed. AI stance: Tests are built AI-first under expert direction, as automation you keep and re-run, not as an agent you ask to poke at the app and hope. ### 06 Release (delivery-flow) Release makes shipping a routine: known contents, known sequence, and a tested way back. AI stance: Releasing a few times a day is now routine for disciplined teams; without release discipline, more output just means more blast radius. ### 07 Monitoring (delivery-flow) Monitoring means you find out first: that the system is up, that the core process inside it works, and whether the outcome strategy intended actually moved. AI stance: Ask the system questions in plain language instead of reading raw dashboards. And never hand an agent broad production access; it can do damage no dashboard ever could. ### 08 Department Management (engineering-os) Department management keeps technology legible and steerable for the business, a department with goals and metrics, not a black box with a budget. AI stance: AI makes the state of your technology legible to non-technical leadership, live metrics and plain-language answers instead of taking anyone’s word for it; the standard it reports against still has to be written down. ### 09 Project Management (engineering-os) Project management makes delivery predictable: the right work, sized to real capacity, finished in the order that matters. AI stance: AI keeps the full project record current - every decision, estimate, and change traceable; what it can’t do is pick the sprint’s one goal, and generated volume is not progress. ### 10 Infrastructure (engineering-os) Infrastructure is what holds when something fails: environments, backups, access, and security, owned by the business rather than the vendor. AI stance: AI does the infrastructure work that used to need a dedicated role - environments, automation, infrastructure as code; access guardrails are what keep an agent from becoming the incident. ## Tools Instruments Brival builds inside client work to enforce its delivery standard — run in production on client projects first, headed for open source. ### TCMS — Test Case Management System URL: https://brival.co/tools/test-case-management Test cases as markdown in your repository, written by coding agents, graded and audited in a web panel. Status: In production use on client projects. Open-source release planned; MCP server, test-automation integration, and CI auto-scaffolding are on the roadmap. ## Case studies ### CtrlCV: From Clickable Prototype to Production MVP in Four Months URL: https://brival.co/case-studies/ctrlcv-academic-cv-mvp Client: CtrlCV · 2026 Stage: Toronto start-up, two self-funding non-technical co-founders, a designed prototype and no engineering Engagement: Fixed-scope, fixed-price MVP across four milestones, kickoff to production handover Stack: Next.js, TypeScript, Supabase, PostgreSQL, CI/CD **We turned a clickable prototype into a production MVP in four months: the full academic CV process, formatted to the letter, at a fixed price.** Two academics, neither of them engineers, had a prototype of CtrlCV and no way to build it. We delivered a production MVP in four months on a fixed price: the profile builder, the publication import, and the CV engine that has to satisfy grant formatting rules exactly. Outcomes: - 40–100 pages — An academic CV runs to 40 or 100 pages and exists in several variants, kept by hand in Word. Grant applications impose exact margins, fonts, section order and citation style, and a document that breaks those rules can be rejected on the formatting alone. [reported by client] - 4 months — Kickoff and architecture, release candidate, testing and UAT, then production rollout to the founders’ own infrastructure. The scope was fixed at the start and nothing was renegotiated mid-flight. [reported by client] [CtrlCV](https://ctrlcv.co) is an academic CV generator: one structured place for a researcher’s publications, grants and mentorship, and an export engine that turns that record into precisely formatted CVs. Its two founders are academics, not engineers. They came to us with a prototype that did everything except produce a CV, and it turned out to be the most useful thing they could have brought. ## The Situation The founders had already done the hard product thinking, and they had something to show for it: a visually complete prototype, built in Replit, that defined the data entry forms and the export configuration page. You could click through the whole product in it. Underneath the screens there was no data model, no persistence and no export engine. That is not a criticism of it. It is what made the engagement work. What they had built was a specification, and it had already settled the questions that normally consume the first month of an MVP: what the product should feel like, what belongs on each screen, what the export step needs to offer. They had validated the concept with trusted colleagues before spending a dollar on engineering. When a scope arrives that clear, a fixed price stops being a gamble for both sides. The step it skipped was the hard one. An academic CV is either exactly right or it is rejected. Researchers keep several variants of a 40 to 100 page document by hand in Word, and grant bodies impose exact margins, fonts, section order and citation style. Formatting is not presentation here, it is correctness. Everything else in the product is scaffolding around that one feature. The constraints were real too. Two self-funding founders, a budget that could not absorb an overrun, a scheduled usability study with students, and no way to evaluate engineering work for themselves. ## What We Did We built the product the prototype specified, on a fixed scope and a fixed price, across four milestones: kickoff and architecture, release candidate, testing and UAT, and production rollout. The profile builder came first, a structured PostgreSQL schema holding the career record, publications, grants, mentorship, behind authentication. Then the CV generation engine, which is where the work actually was: PDF and DOCX export driven by strict formatting rules, APA and MLA citation styles, and the export controls the founders had designed, filter by date range, select or deselect individual records and sections, reorder sections. We also integrated ORCID and publication import by DOI, PMID, BibTeX and RIS. This one matters more than it looks. A researcher with several hundred publications will not retype them, so without import the profile builder is a chore nobody finishes and the rest of the product never gets used. The project ran with the founders rather than around them: weekly progress and review calls, email in between, and direct access to the project manager and the lead developer instead of an account layer. Anything out of scope or at risk of not landing was named when it came up, not at the end. Because neither founder was technical, we also walked them through configuring the third party services the product depends on, so the platform ended up in their accounts, under their control. ## What Shipped A working MVP, deployed to the founders’ own infrastructure, handed over completely: the Git repository, the database schema and migrations, deployment documentation, and CI/CD configured. They own the product and the infrastructure it runs on. There is nothing of ours they need to keep paying for to keep it running. It was delivered in the window we agreed, March to June 2026, with the scope we agreed. In the client’s published review, Brival scored 5.0 on quality, schedule and cost, with no areas for improvement identified. ## What We Cannot Claim Yet No business outcomes yet. As this is written, handover is complete and CtrlCV is in private beta, which means the product is in real hands but has not yet produced the numbers that would matter: adoption, hours saved per CV, what the usability study finds. None of that exists, so none of it appears above. Every number on this page is either a delivery fact or the client’s own description of their problem, and each one is tagged as such. Those numbers come out of the private beta, and we will publish them here when there are some. Until then the honest version of this case study is the one you just read: a scope made precise by a good prototype, built for the price and on the date we said, and handed over whole. ### Healthcare Platform: Modernization Under an Enterprise Security Deadline URL: https://brival.co/case-studies/healthcare-platform-modernization Client: Healthcare start-up · 2022– Stage: Healthcare start-up at product-market fit; aging codebase now facing enterprise security reviews and an end-of-life stack Engagement: Project development (original platform) → incremental modernization, delivered while live Stack: Django, Docker, CI/CD, REST API **When growth raised the bar, with enterprise security reviews and a stack against its end-of-life clock, the fix was hardening, not a rewrite: the same live platform taken from zero automated tests to 95%+ coverage on its core logic and brought to a security-review-ready posture, modernized while it kept running.** A healthcare start-up’s platform had reached product-market fit on an aging codebase. Then enterprise clients’ security reviews and an end-of-life stack raised the bar. We modernized the live system incrementally, without interrupting operations: MFA and a full audit trail, off that stack onto supported versions, and a regression test suite where there had been none. Outcomes: - 0 tests → 95%+ — A system carrying real clinical load had effectively no automated tests. We brought core business logic and the views that carry it to 95%+ coverage each: the regression safety net a platform handling medical data needs before anyone changes it under pressure. [measured] - failed → cleared — Now clears the enterprise security reviews it couldn’t before: MFA, 30-day credential rotation, and a full audit trail. [practice] - 100s GB / month — The modernization ran incrementally on the live platform, never a clean-slate rewrite, while it kept serving thousands of inquiries and hundreds of gigabytes of medical data a month. Operations never stopped. [reported by client] A healthcare start-up’s platform had reached product-market fit, and growth had moved the goalposts. Enterprise clients now ran vendor security reviews the platform couldn’t yet pass, and its stack was approaching end-of-life. We modernized it without ever stopping the business: from zero automated tests to 95%+ coverage on the code that runs it, onto a supported stack, with MFA and a full audit trail. We name the metrics, not the client. ## The Situation Our client is a healthcare start-up whose web application runs the operation: it processes sensitive medical data, with operational automation, system-to-system integrations, and two-factor authentication on access. We had built the original platform, and it worked: it carried the product all the way to product-market fit, handling thousands of inquiries and hundreds of gigabytes of medical data a month. Then growth changed what it had to be. Winning and keeping enterprise clients meant clearing their vendor security due diligence, and the earlier posture wasn’t built for that scrutiny. The stack itself was nearing its end-of-life cutoff, a hard support deadline rather than a preference, with security and maintenance debt compounding behind it. Access controls sat below what a platform handling medical data needs to clear such a review. And underneath all of it, a system carrying real clinical load had effectively no automated tests, on a codebase shaped to reach a market, not to be built on. We were brought back to modernize the live platform against that deadline, without stopping operations. ## What We Did The constraint set the method: modernize incrementally while the platform kept running, never a clean-slate rewrite. The hard part of a job like this is the ramp-up: re-learning a codebase a couple of years old, built for speed rather than for the next team, before a single change can be made safely. AI-augmented delivery is what made that painless. It accelerated reading and mapping the existing system, drafting the test suite that had never been written, and working through the dependency upgrades, keeping the catch-up economical against a fixed deadline instead of letting it become the bottleneck. We hardened the platform for the security reviews: multi-factor authentication, password rotation enforced every 30 days, and a full audit trail across the system, the access-control and traceability evidence a vendor security review asks for on medical data. We took the platform off its end-of-life stack, upgrading everything to supported long-term-support versions and dropping unused components, which cleared the deadline and shrank the maintenance and security surface at once. We built the regression safety net that wasn’t there, bringing core business logic and its views to 95%+ coverage each, from none. We Dockerized the environment and set up automated CI/CD, turning releases into a repeatable path. And we exposed a clean reporting API for the enterprise systems to consume results, leaving the platform more open than we found it. ## What Moved Underneath, the test numbers: 82% coverage across the repository and 95%+ on the business logic and views, from a starting point of none. The security posture is what the enterprise vendor reviews were asking for: MFA, 30-day credential rotation, and a full audit trail. The end-of-life deadline is cleared, the platform now on supported versions with dead components removed, and releases run through Docker and automated CI/CD. The scale figures, thousands of inquiries and hundreds of gigabytes a month, are reported by the client, not measured by us; that the modernization landed without interrupting any of it is the part we deliver. ## Where It Went None of it was a rewrite. The platform that won product-market fit is the same one now cleared for enterprise scrutiny: hardened in place, against the deadline, while it kept running. ### TripHero: QHero Platform Rebuild URL: https://brival.co/case-studies/triphero-qhero-rebuild Client: TripHero · 2024– Stage: Logistics startup post-PMF; goal of growing from 6 to 30 hotels in two quarters Engagement: Codebase audit → fixed-price rebuild → ongoing development incl. certification support Stack: Directus, Node, React + Material UI, Google Cloud OCR, ShipStation, Shopify **A proof of concept that had outgrown itself was rebuilt to a production standard in under 12 weeks at a fixed price, and the same discipline carried it through security certification into an international hotel network.** QHero automates hotel package handling. The platform meant to grow a logistics startup from 6 hotels to 30 in two quarters was proven in production but still on its proof-of-concept build, so we rebuilt it for the standard that growth required. Outcomes: - <12 weeks — The rebuilt web application shipped within the expected budget and hit the quarter’s KPI of an operational platform. [measured — historical] - SOC 2 — After several months of certification work the product was approved for use in an international hotel network. [measured — historical] - 2+ hrs/day — QHero’s own product marketing reports hotel staff saving 2+ hours a day and package-handling revenue up by as much as 25%. [reported by client] QHero, TripHero’s hotel package-and-shipping logistics platform, had proven itself in production but was still running on its proof-of-concept build. We audited the codebase, rebuilt the platform as a production web application in under 12 weeks within the expected budget, and then supported it through the security certification that opened an international hotel network. ## The Situation TripHero helps hotels and individuals transport travelers’ luggage; QHero handles hotel package logistics from receipt to delivery: OCR scanning, automated tracking. The product began as an internal proof of concept and, after a year-plus of live testing with real hotels, had established product–market fit. The company set a goal of growing from 6 to 30 hotels within two quarters. The PoC couldn’t carry that: the interface lagged, the platform slowed visibly under even a small client base, hotel staff copied data by hand between third-party tools, and features kept breaking. One more door was locked outright: an international hotel network required a security certification. ## What We Did We audited the existing codebase first. The recommendation was a rebuild on a stack aligned with the long-term perspective. We don’t lead with rebuilds: at PoC scale, rebuilding was cheaper than untangling temporary-solution choices made under real client load, and that math flips on larger systems. The rebuild ran as a bounded project: a production web application in under 12 weeks within the expected budget. We built a Hotel Dashboard for daily operations (guests, shipments, groups; tablet-optimized for the front desk) and an Admin Dashboard for full oversight, and automated the most time-consuming steps: AI-based OCR on Google Cloud reading shipping-label and guest information, automated label generation through ShipStation, checkout via Shopify. Then ongoing development: onboarding new hotels, market-driven features, and hands-on support through the security certification, hardening the solution, putting protocols in place, providing technical insight on the setup. Our side was deliberately small: a PM, full-stack developers on React and Node, architect support, and QA. ## What Moved The production platform shipped in under 12 weeks on the expected budget, meeting the quarter’s operational KPI. After several months of certification work, the product passed and is SOC 2-certified, approved for use in an international hotel network. The end-user numbers belong to the client: QHero’s own product marketing reports staff saving 2+ hours a day and package-handling revenue up to +25%, their figures from their own product marketing, not our instrumentation. ## Where It Went Ongoing development continues alongside the client’s CEO and product manager: new hotels onboarded, features driven by the market. The shape repeats a pattern we see often: PMF established on a build that can’t carry the growth. The fix wasn’t heroics; it was an audit, a clear recommendation, and a bounded rebuild. ### Ad-Tracking Platform: From Crashing Every Few Days to Near-Zero Incidents URL: https://brival.co/case-studies/ad-platform-modernization Client: Performance-marketing company · 2019– Stage: Revenue-critical software, every campaign click flows through it Engagement: Original build → infrastructure modernization → ongoing maintenance, plus team augmentation on other supporting initiatives Stack: Django, Docker, PostgreSQL, CI/CD, DataDog, Sentry **When traffic growth started crashing a revenue-critical ad platform every few days, the fix was infrastructure, not a rewrite.** A performance-marketing company's ad platform, two Django systems tracking every click its campaigns buy, was crashing every few days as traffic grew. We modernized the infrastructure around the code, not the code: Docker, managed PostgreSQL, automated deployments, observability, a live production migration. Near-zero incidents since, at millions of requests a month. Outcomes: - 10M reqs/month — The platform serves millions of requests a month, approaching 10 million in peak season, roughly 2 million tracked clicks a month. Seasonal peaks run at several times the quiet months, absorbed without any scaling intervention. [measured] - 1+/week → <1/quarter — Before the modernization the system crashed every few days, and the crashes clustered where they cost the most: during the highest-load campaigns. Since the migration it runs unattended for months at a time. [measured — historical] A performance-marketing company runs its campaigns through an ad platform we built: two Django systems that redirect and track every click across more than a dozen domains. As traffic grew, the platform began crashing every few days, worst during the highest-load campaigns. We modernized the infrastructure around the application, not the application itself, and migrated it to managed cloud. It has run at **near-zero incidents** since, serving **millions of requests a month**. We name the metrics, not the client. ## The Situation Our client operates in performance marketing: affiliate campaigns whose every click is money, routed through a platform of two Django systems. One serves the public campaign links across more than a dozen domains; the other validates and records each click, running campaign checks, geo rules, click and lead caps, per-IP rate limits, and link rotation before anything gets counted. When this platform is down, campaigns keep spending while the clicks that justify the spend go unrecorded. The platform had been built years earlier, for the traffic it had then, and the application code was doing its job. The environment underneath was the problem: a single bare-metal server, a self-managed database on the same box, storage backing up, deployments done by hand, and no observability to see any of it coming. As campaign traffic grew, the system began crashing every few days, and the crashes clustered exactly where they cost the most: during the highest-load campaigns. ## What We Did Two goals, set with the client: improve product stability, and make the stack modern, secure and maintainable. A rewrite was never on the table; the application logic had proven itself for years. The scope was everything around it. We containerized both applications with Docker and moved the operation to cloud infrastructure: a managed PostgreSQL database with automated backups, a dedicated web server with reserved networking, and deployments automated from the repository instead of run by hand. We integrated observability, monitoring and error tracking (DataDog, Sentry), with alerts on the failure modes that used to be invisible: resource exhaustion, disk filling up, application errors. And we wrote the operational documentation, an infrastructure map and runbooks for deploys, restarts, and log review, so the platform's operation no longer lived in anyone's head. The migration itself: DNS TTLs lowered a day ahead, the database copied to the managed instance, traffic switched, the old server retired. A year later we returned for a database performance pass, because reporting at this scale is a balancing act: every index that speeds up a real-time report slows down the writes underneath it, and this platform writes on every click. We adjusted the data structures and applied indexing incrementally, measuring at every step which changes were hurting database writes in our case, and settled where the compromise holds: reports close to real time, on a database that still keeps up with the click stream. The reporting queries themselves were rewritten to cut the load the daily reports put on the database. The relationship has continued since: we maintain the platform and have augmented the client's team on other supporting initiatives. ## What Moved Stability first, because it was the goal that triggered the work. Before the modernization, crashes every few days under campaign load. Since, near-zero incidents: the platform **runs unattended for months at a time**, at a load average around five percent of its capacity. Then the scale that stability holds up under. Measured from 20 months of production access logs: **5 to 6 million requests a month** on average, **9.8 million in the peak month**, peak days around **640,000 requests**, roughly **2 million tracked clicks a month**. Seasonal campaign peaks run at several times the quiet months, and the platform absorbs them without any scaling work. All of these figures are measured by us from production logs. ## Where It Went The codebase that was crashing is the same codebase serving the peaks now. What changed is everything around it: the platform under the software, the automation around releases, and the visibility that turns surprises into alerts. When a working system outgrows its infrastructure, that is the honest scope of the fix: modernize the platform under it, and leave proven software alone. ### Truity: Personality Testing Mobile App URL: https://brival.co/case-studies/truity-personality-app Client: Truity · 2023– Stage: INC 5000 company; one-time purchase model capping customer LTV Engagement: Fixed-price MVP → ongoing development under a monthly budget cap Stack: React Native, Node, Directus, PostgreSQL **A business experiment was priced like one: a fixed-price, twelve-week MVP let an INC 5000 company test a new revenue model without betting the roadmap.** An INC 5000 personality-testing company hit a ceiling: customers bought once. We delivered a subscription mobile app to lift lifetime value, then grew it into multi-year development and a full platform migration. Outcomes: - <12 weeks — A working subscription app shipped fast enough and cheap enough that testing a new revenue model never threatened the core business. [measured — historical] - Multi-year — Once the model was proven, the MVP became the foundation: multi-year development and a backend migration that modernized the whole platform, not just the experiment. [practice] Truity sells personality tests that millions of people have taken; the growth obstacle was customers purchasing once. We delivered a subscription mobile app, with daily mood tracking, personality surveys, progressively unlocked insights, in under 12 weeks on a fixed price, then carried it through years of ongoing development, including a full backend migration. ## The Situation Truity, a California company that made the INC 5000 fastest-growing list three years running, had a large website funnel and customers who nearly all purchased once. A mobile app was the play to convert that funnel into a subscription product and raise customer LTV. The constraints were the real brief: fast, at bounded cost, and without distracting the core business that was funding the bet. ## What We Did Fixed-price MVP first: a working React Native cross-platform app delivered in under 12 weeks within a deliberately small fixed budget. What that version deliberately didn’t have: a backend the team would enjoy operating. We ran it on Firebase managed through Retool: workable, not pleasant, and chosen because operability wasn’t the question the MVP existed to answer. Then the engagement moved to ongoing development under a monthly budget cap, working closely with the CEO and the in-house CTO and product designer who led product direction. When operability did become the question, we migrated the backend from Firebase + Retool to Directus on Node and PostgreSQL, giving the team a structured, responsive content-management interface. We kept iterating: push notifications, convenient authentication, a GenAI chatbot integration, and codebase updates as the platform aged. QA was part of the team throughout, so sustained change didn’t erode quality. ## What Moved The MVP shipped in under 12 weeks on its fixed budget: the side bet validated at strictly bounded cost, which was the metric the engagement was designed around. The app is live on Google Play and has been iterated continuously since; that is a fact of record, not a performance metric. The Directus migration gave Truity’s team a content interface they could actually run the product with, a qualitative outcome, as the team reported it. What we don’t publish: subscription-revenue figures. That instrumentation belongs to the client, and we don’t claim numbers we didn’t measure. ## Where It Went The engagement grew from a small fixed-price MVP to multi-year ongoing development with the same team. That progression, from bounded experiment to sustained product work, is the pattern we now design engagements around: price the experiment like an experiment, and only scale the spend once the business question has an answer. ### SNE Częstochowa: Moving a Course Operation Into Software in Seven Weeks URL: https://brival.co/case-studies/sne-czestochowa-course-platform Client: SNE Częstochowa · 2025– Stage: Catholic evangelization community: hundreds of members in dozens of volunteer groups across ten local schools; a course calendar run on phone calls, web forms and bank statements Engagement: Fixed-price projects: the community app first, then a seven-week rebuild with the course operation built in Stack: Django, React Native, Expo, PostgreSQL, Celery, Tpay, Twilio **Seven weeks to move a community’s courses off the director’s phone: the app rebuilt on the backend that already worked, with sign-up, payment, crew and settlement running through it.** SNE Częstochowa runs retreat courses for hundreds of participants a year, organized by volunteers across ten local communities. Sign-ups came through an old website form, payments were matched against bank statements by hand, and each course ran on the director’s phone calls. We rewrote their app natively, kept the backend, and put every step of a course into it. Outcomes: - 7 weeks — Seven weeks from the decision to rewrite the app to the new version in the stores. Online payments switched on two days later and took their first payments within a day. [measured] - 66 → 799 — The backend was kept and extended rather than rewritten: every original data model is still in place, and the automated tests grew twelvefold alongside the new features. [measured] - 100s a year — Hundreds of members in dozens of volunteer groups, and hundreds of course participants a year, now run through one system: the season’s courses, their sign-ups, payments and settlements. [reported by client] SNE Częstochowa is a Catholic evangelization community that has run retreat courses and events since 2011, across ten local schools, by volunteers. We built its first app, then rebuilt it in seven weeks and moved the organization’s courses into it, from the sign-up to the settlement. The app is what people see. The process is what changed.
Payments that match themselves
Online payments land on the participant’s record automatically, instead of someone matching bank statements to course lists.
Sign-up without a coordinator
People sign up in the app or from the website, answer the course’s own questions, and get every follow-up message on schedule.
One view of the organization
Every course’s stage, every payment and every member’s history in one place, instead of on the director’s phone.
Seven weeks to the stores
A native rewrite of the app, from the decision to rebuild to the App Store and Google Play.
Payments from day one
Online payments switched on two days after the release and took their first payments within a day.
Nothing thrown away
Every original data model kept and extended, with automated tests grown from 66 to 799.
Course details open to all participants, and participation management layer for manager.
Sign-up and payment now happen in one place. The community’s website lists the season’s courses straight from the platform and hands the sign-up to it, with each course’s own questions: dietary needs, parish, previous courses. Payment runs online through Tpay and is matched to the participant automatically, with a bank transfer as the fallback. Crew members owe nothing, as a rule in the system rather than a note to the treasurer. Members apply to serve on a course from the app, and the director approves or declines. Every new course starts with a default schedule of six messages, from the sign-up confirmation to the note after the course, plus messages to all participants or the crew by email or SMS. Each member carries one history: the courses they attended or served on, their service roles, and the groups they belonged to over time.A settlement that will not submit until attendance and the required summary answers are in, and a message to everyone on the course.
## What Moved The organization’s courses now run through one system instead of forms, mailboxes, bank statements and phone calls: - **Sign-up**: from a web form and a shared mailbox to a self-service sign-up, in the app or from the website, with the course’s own questions. - **Payment**: from bank statements matched by hand to online payments matched automatically. They took their first payments within a day of switching on. - **Crew**: from a separate form and phone calls to a request in the app that the director approves. - **Messages**: from emails written per course to a schedule every course starts with. - **Settlement**: from an ad hoc reckoning per course to a stage every course enters when it ends, which will not close with anything missing. - **Oversight**: from one person’s phone to every course’s stage visible in the app. Underneath, the season’s 19 courses sit on a backend that was extended, not replaced, with its automated tests grown from 66 to 799 and checks that run on every change. ### Evolve Edits: Web Platform Rebuild URL: https://brival.co/case-studies/evolve-edits-platform-rebuild Client: Evolve Edits · 2022–2023 Stage: Photo post-production company; decade-old portal moving terabytes a month through one on-premise machine Engagement: Consulting → fixed-price rebuild, canary rollout → ongoing maintenance & monitoring Stack: Django, FastAPI, React, DigitalOcean, Podio, Authorize.net **A decade-old, revenue-critical platform was modernized while the business kept operating: 90% fewer downtime incidents and 80% lower infrastructure cost.** A photo post-production company’s decade-old client portal funneled terabytes a month through a single on-premise machine. We rebuilt it as a cloud platform and released it by canary rollout, no big-bang cutover. Outcomes: - 12 incidents — Incidents collapsed once the revenue-critical upload flow stopped depending on a single on-premise machine. [measured — historical] - ~30 MB/s → uncapped — Uploads ran through one machine, capped and shared across every client. Direct-to-object-storage uploads let each client upload in parallel at their own bandwidth, concurrently. [measured — historical] - $1,700 → $300 (−80%) — Provisioned storage priced for peak was replaced with object storage; costs fell while the platform stored more data for longer. [measured — historical] Evolve Edits’ decade-old WordPress client portal moved terabytes of photos a month through a single Mac Mini at company HQ. We rebuilt it as a cloud platform on a fixed price and released it by canary rollout while the business kept operating. Downtime incidents fell 90%; infrastructure cost fell 80%. ## The Situation Evolve Edits is a post-production company for photography: photographers worldwide send their photos in for editing, with software sold as part of the service. Their client portal had grown for over a decade inside a WordPress marketing site fueled by outdated plugins. Files uploaded by clients went directly to a Mac Mini at company HQ, which suffered frequent downtime and kept running out of storage while the business moved terabytes per month. Uploads were throttled by one machine’s network throughput and disk write speed; file transfers ran over FTP with no alerts, synchronization, or automation, though international clients needed 24-hour self-service. And the AWS storage that backed it was provisioned for peak and paid for always. ## What We Did Consulting first: we assessed the existing solution, identified the improvement areas, and presented the project plan. Then a fixed-price rebuild, delivered over roughly nine months. We built a React + Material UI client portal, a Django backend, and a FastAPI file-transfer microservice. The decision that mattered most was architectural: direct-to-object-storage uploads, which removed the single-instance bottleneck entirely instead of buying a bigger instance. We migrated from AWS to DigitalOcean, replacing provisioned storage with object storage, integrated two-way with Podio via webhooks on task changes, exposed data to Looker Studio and automation platforms, and wired Authorize.net payments. Release was a canary rollout, gradual exposure to a small subset of users, because the portal is a revenue-critical flow and a big-bang cutover would have bet the entire client base on day one. That made the release slower than it could have been, deliberately. ## What Moved Downtime incidents fell from 125 to 12 per quarter, a −90% change, after the on-premise dependency was removed. The comparison is the same quarter a year apart, the baseline being the quarter the new platform shipped in, and both figures are as recorded at the time, not products of the standardized definitions our current proof loop uses. Infrastructure cost fell from $1,700 to $300 per month, −80%, while the platform stored more data for longer. Clients now upload in parallel at their own bandwidth and reported the difference unprompted; we record that as a qualitative observation, not a number. The human element left the transfer pipeline: Podio integration and automated workflows replaced manual FTP handling. ## Where It Went After launch the engagement transitioned to ongoing maintenance and monitoring, with our PM and tech lead working directly with their production manager and in-house technical role. ### Sensitive-Data Platform: Mobile App Rebuild and Enterprise Product URL: https://brival.co/case-studies/personal-safety-app-rebuild Client: Consumer platform, sensitive data · 2021– Stage: U.S. consumer company handling sensitive personal data; investor-promised B2B milestone blocked by the codebase Engagement: Audit → consulting → fixed-price rebuild → ongoing development Stack: Firebase, Directus, Node, React, native iOS, native Android **The cheap truth comes first: a small fixed-price audit re-sequenced a company roadmap, and the stabilize-before-build order it dictated is why the investor-promised product shipped at all.** A consumer company whose product holds sensitive personal data and runs on a chain of third-party integrations needed an investor-promised B2B product its existing foundation wasn’t built to carry. We rebuilt the platform in 7 months, and the enterprise product shipped on top. Outcomes: - 7 months — The secure, scalable replacement platform was delivered in 7 months, and the B2B enterprise product promised to investors shipped on top of it. [measured — historical] - ~1% — A small fixed-price audit, about one percent of what the engagement eventually became, converted “uncertainty about what we own” into a risk-ranked picture the board could act on. [practice] Our client runs a consumer platform that holds sensitive personal data and moves it through a chain of third-party integrations. A small fixed-price codebase audit showed the existing foundation wasn’t built to carry their investor-promised B2B product, so the roadmap was re-sequenced: we rebuilt the platform in 7 months at a fixed price, and the enterprise product then shipped on top of it. We name the metrics, not the client. ## The Situation The client is a U.S. consumer company running a real-time platform that stores sensitive personal data and moves it through a chain of third-party integrations, every one a dependency that has to hold the moment it’s needed. They came to us to take over development and to build a B2B enterprise product promised to investors as a major roadmap milestone. The existing codebase, a Firebase mobile backend, a PHP web portal, and native iOS and Android apps from a prior team, had been built to reach the market, not to carry an enterprise product or the security and scale that came with it. The board wanted ground truth about the asset before committing to the new initiative. ## What We Did The sequence our service ladder is now built around, executed here first. A small fixed-price codebase audit across all platforms mapped what would and wouldn’t carry the enterprise build: code quality, maintainability, security, scale. Consulting turned the findings into a plan, covering security, performance, tech-debt catch-up, and the board made an informed decision: the enterprise product was deliberately postponed until the foundation was ready to carry it. Telling a client to delay the product they had promised investors is not what anyone hoped to hear from us; it is also the reason that product shipped at all. The rebuild then ran as a fixed-price project, with the delivery risk on us at the riskiest step: re-designed architecture, fixed API communication, new apps on a modern stack. The hard engineering lived in the product’s nature: sensitive personal data in motion, real-time behavior, and a chain of third-party integrations, every one a point that has to hold the moment it’s needed. ## What Moved The rebuild was delivered in 7 months at a fixed price. The audit converted uncertainty into a risk-ranked picture and the roadmap was re-sequenced on evidence rather than optimism. The enterprise product, the investor milestone, shipped on top of the rebuilt foundation; that is a fact of record, not a performance metric. These figures are as recorded at the time; they pre-date the standardized metric definitions our current proof loop uses. ## Where It Went The engagement continues as ongoing development under a monthly budget, building out the enterprise product and iterating on the platform alongside the client’s CTO and a dedicated product manager. The rebuilt platform has since passed a security certification. Diagnosis, then stabilization, then the new initiative, with a fixed price at the risky step: this engagement is where that order proved itself. ### Plunjr: MVP and Platform Evolution URL: https://brival.co/case-studies/plunjr-mvp Client: Plunjr · 2020– Stage: U.S. startup; app live for a year, 100+ installs a day, very few real customer calls Engagement: Fixed-budget MVP → ongoing development through several pivots Stack: Django, React, native iOS, native Android, Twilio Video / Zoom, Stripe, HomeDepot APIs **Quality standard is a dial, not a default: a deliberately raw, fixed-budget MVP answered the market question, and the same team turned the codebase production-grade the moment paying customers arrived.** A home-services startup had a feature-packed app live for a year and almost no real calls. We stripped the question down to a single Call Now button to find out why, then evolved the product through several pivots. Outcomes: - 1 button — A single Call Now button isolated the one question that mattered, will customers actually call, and surfaced real motivations and objections before any further feature spend. [practice] - $10K — The prior feature-rich app had cost many times this and kept billing to extend features it had never validated. The MVP answered the one question the business depended on first, before another dollar went in. [measured — historical] Plunjr connects homeowners with licensed plumbers over video calls. We built them a deliberately raw, fixed-budget MVP, a single Call Now button, that answered the question their feature-packed app hadn’t: will customers actually call. We then developed the platform through several pivots into a production-grade property-management product. ## The Situation When we engaged, Plunjr already had a production app live for over a year, with video calls, payments, scheduling, 100+ installs a day, yet very few real customer calls. The app had been built by a U.S. agency at a cost the company couldn’t sustain, and video streaming ran on a niche early-stage framework, below the reliability bar Skype and Teams had already set in users’ heads. The system had plenty of features and no answer to the only question the business depended on. ## What We Did The first engagement was a fixed-budget MVP, built in a few weeks, codenamed “The Big Button Project.” One giant Call Now button. It was deliberately raw: no new features, no polish, an experiment rather than a product. Its only job was to isolate the primary thesis and surface real customer motivations and objections. An MVP this raw buys you an answer, not an app. That trade was the point. From there, ongoing development iterated on market feedback through several pivots. We migrated video to industry-standard platforms, Twilio Video, then Zoom, with native calling via CallKit, and expanded to Android and web interfaces for technicians and property managers. When the first paying clients arrived, we turned the quality dial deliberately: test coverage, automation, refactoring, infrastructure sized for load, documentation catch-up. The rapid-experiment codebase became production-grade because paying customers made that the right standard, not earlier and not later. ## What Moved The market question was answered for a few weeks of fixed-budget work: client motivations and objections were identified before further feature spend. After the platform migration, video and calling reliability matched the leading communicators, and daily video calls across iOS, Android, and web ran undisrupted once paying clients depended on it; these are qualitative observations from the time, not numbers we recorded. The contrast that mattered was cost-to-learn: the MVP answered the market question for roughly $10K, a fraction of what the prior feature-rich app had cost, with specialists flexed per stage. ## Where It Went Over the engagement’s multi-year arc the product grew into a property-management platform: multi-party real-time chat, payments, e-commerce, nested organizational permissions, Stripe and HomeDepot integrations. Our team worked alongside the founder, a product manager, and supporting marketing and analytics roles. The lesson we kept: calibrate the quality standard to the product’s stage. Gold-plating an experiment wastes the budget, and shipping an experiment to paying customers wastes the trust. ### FanServ: NBA Team Fan App Rebuild URL: https://brival.co/case-studies/fanserv-nba-fan-apps Client: FanServ · 2018 Stage: U.S. sports & entertainment agency; official NBA team app serving hundreds of thousands of fans Engagement: Full technical ownership: development, PM, QA · one focused engagement, ~2,500 team hours Stack: Django, AWS, Celery, Angular, native iOS, native Android, Firebase, APNS2, Node **When failure is public and the load is spiky, discipline is what holds: the rebuilt platform pushed to tens of thousands of devices in under half a minute and kept game state 12× fresher.** The official app of an NBA team, serving hundreds of thousands of fans on game night: we took full technical ownership of the backend, native apps, push pipeline, live feed, and CMS for the agency that owned the product. Outcomes: - 60s → 5s (12×) — In-app game state went from a minute behind the actual game to five seconds, delivered to devices in real time over WebSockets instead of a 20-second pull. [measured — historical] - 20–30 seconds — Post-quarter notifications during live games reached every registered device in 20–30 seconds. Previously, some fans got updates late or not at all. [measured — historical] We rebuilt the technical platform behind an official NBA team’s fan app for FanServ, the agency that owned the product: backend, native iOS and Android apps, push pipeline, live feed, and CMS. Live feed refresh went from 60 seconds to 5; push notifications reached tens of thousands of devices in 20–30 seconds during games. ## The Situation FanServ, a U.S. digital agency for the sports and entertainment industry, was tasked with redesigning the official app of an NBA team whose app served hundreds of thousands of fans. The existing app ran on outdated core components. Push delivery was unpredictable: some fans got game updates late or not at all, at tens-of-thousands-of-devices scale. The in-app feed lagged the actual game: the league’s legacy FTP/XML pipeline and a 60-second refresh couldn’t keep up. Editors managed content through a raw Django Admin, typing in by hand what league APIs could have automated. FanServ’s product owner and designer led the vision; we owned the complete technical execution: development, project management, and QA. ## What We Did We rebuilt the platform from the ground up under live-event constraints: hundreds of requests per second during games, with any failure visible to tens of thousands of fans in real time. The push pipeline moved to parallel processing across multiple Celery queues with multi-threading, and delivery migrated to APNS2 and Firebase. For the live feed, we replaced the league’s legacy FTP/XML processing with official-API scraping at a tighter interval, and replaced the mobile data pull with WebSockets so devices receive game state as it changes. Around that core: new native iOS and Android apps with a new design and fan-requested features, a new Django backend on AWS built from scratch, and a dedicated Angular management portal replacing the raw Django Admin. One constraint shaped the tail of the project: no staging environment reproduces a real game-night spike, so we treated the first live games as part of the engagement and monitored them end to end. ## What Moved Feed refresh went from 60 seconds to 5, and delivery to devices moved from a 20-second pull to real-time WebSockets. Game state in the app was 12× fresher. Post-quarter push notifications during live games reached all of tens of thousands of devices in 20–30 seconds. Fan engagement improved with the redesigned app and editors were happier with the dedicated CMS; those are qualitative observations from the time, not numbers we recorded. ## Insights ### Building in Healthcare or Finance: What You Can’t Add Later (HIPAA, SOC 2, GDPR) URL: https://brival.co/insights/2026/09/building-regulated-software Category: Methodology · By Mat Kupczyk · 2026-09-10 Most compliance work in a regulated product can be added whenever you get to it: multi-factor login, background checks, policies, scanning. A smaller part cannot, because it is about the past. Records that are never overwritten, a trail of every access, field-level encryption and separation between clients only count if they were running when the data arrived. It usually starts with a sentence that sounds harmless: it is just a form. On screen, it is. Someone types in some details, someone else signs them off, and anyone who has moved a process from paper to digital has done it before. The video above is the full walkthrough. This is the written version, with the four things to do this month at the end. ## What the Audit Asks That familiarity is the trap. It makes the product look like any other business process: map it out, build the form, done in two weeks. Eighteen months later, someone sits down to audit you, and they do not ask about the form. They ask for three things: - Who opened this record in March. - Whether anything in it changed after it was signed. - Proof, from outside your own systems, that you were scanning for vulnerabilities and patching them, month after month. None of those is a feature you forgot. Each one is a decision you made, or did not make, before you started building. ## What Can Wait, and What CannotSome compliance work you can always catch up on. Start it next month or next year and you will be fine. Some of it you never can, because it is about the past: it only counts if it was already running when the thing happened.
You can put a lock on the door tomorrow. You cannot write down who came in yesterday.
“We don’t know” is what fails an audit. Not a weak control. No answer at all.
### Encrypt the Fields Everyone plans for encryption in transit and at rest. The part teams skip is individual columns: the name, the ID number, the diagnosis, encrypted as fields, so that reading the database is not the same as reading the data. The same applies on the device if your app works offline. You can add this later: migrate, re-encrypt, and absorb a painful project. What you cannot fix is the past. Every backup taken before the migration still holds the data in the clear, and so does every log line where a developer printed a record while chasing a bug. Encryption fixes today forward. The best you can do for the past is keep it in cold storage and let a retention policy age it out. ### Separate Every Client This is the expensive one. Each client’s data lives in its own space with nothing crossing, which means every query, every join, every background job and every report carries that boundary. Building for one client first is easy. Adding the second one later is not a feature, it is a breaking change: you edit every path in the system that touches data, and hope you found them all. And when you are done, you still cannot prove the boundary held before you built it. ## Developers Without Real Data On most regulated engagements, your developers will not be allowed near real data, and that changes how every bug gets fixed. Usually it is the client’s contract that says so: the data does not leave the region, and often it does not leave their environment at all. People assume this is HIPAA. It is not, because HIPAA has no residency rule. Other rules do exist around it. Since April 2025, a US Department of Justice rule has restricted bulk transfers of Americans’ health data to a short list of countries of concern, and banned some kinds outright. Florida requires providers using certified electronic health record systems to keep patient data in the continental US, its territories or Canada. But mostly it is the contract, and the contract is usually stricter than any of them. So when production breaks, the engineer does not pull a copy and debug it. They work through a VPN and a screen share with someone who is allowed to look, or against a made-up dataset that has the shape of the real thing and none of the content. Budget for it. It is not a line in a policy, it is a tax on every bug you will ever fix. And the made-up dataset, or the anonymization job behind it, is real software that someone has to build and keep current. ## Automation Needs a Name In a regulated system every action carries a person’s name, and that quietly caps what you can automate. Who approved this. Who released this. Who exported this. Not an account, a person. About a year in, someone says the obvious thing: this step is slow and boring, automate it. Or better, have the AI do it. Then you hit the wall. Put a service account on the action and an auditor sees a robot doing work that needed judgment, and asks who authorized each one. Put a person’s name on it and that person is now answerable for something they never did and never saw. The fix is to record three things instead of one: ``` actor what did it: a person, a scheduled job, an agent on_behalf_of which person it acted for authorization the standing permission that allowed it ``` That is a schema decision. If your audit table has a single user column, your automation roadmap is already capped, and nobody has told you yet. It matters now because everyone wants agents in the workflow. In a regulated space, the limit is not what the model can do. It is whether your system can say whose decision it was. ## What to Do This Month Four things, and all of them fit inside a month. ### 1. Write Down Your Rules One page, listing what you are actually under. For healthcare in the US that is HIPAA, and NIST publishes guidance on implementing it (SP 800-66). In Europe it is GDPR, plus whatever your own country adds for health data. Check two more things while you are there. First, if your software influences a clinical decision, it may be regulated as a medical device: a separate approval process on a different timeline, and not something to discover in year two. Second, if you are under HIPAA, much of what is described above sits in the rules as “addressable,” which people read as optional. An update proposed in January 2025 would make most of it mandatory. As of September 2026 it is still a proposal, pushed to 2027. Build it now anyway. It is the cheap version of a change that is coming. ### 2. Start the SOC 2 Clock If you will ever sell to enterprises, treat SOC 2 as a day-one requirement. They will ask for it, or for a security review that covers the same ground. When a healthcare platform we built started winning enterprise clients, their reviews asked for multi-factor login, credential rotation and a full audit trail, and [the platform had to be hardened before it could pass](/case-studies/healthcare-platform-modernization). There are two kinds of SOC 2 report. Type I says your controls exist on a given day. Type II says they worked across a period, usually around six months for a first report, and it is the one enterprises want. You cannot move those dates backwards, so the clock starts the day you switch the controls on, not the day you decide you want the report. It is the logbook again: you cannot catch up on a period nobody was watching. ### 3. Insure Both Sides Whoever builds the product should carry professional indemnity, for getting the work wrong, and cyber liability, for a breach. Underneath the insurance sits a contract: a Business Associate Agreement (BAA) in the US if they touch patient data, a data processing agreement in Europe. The contract makes them liable. The insurance makes that liability worth something, because a liability clause against an uninsured company is a piece of paper. Get your own cover too. The regulator comes to you first, whoever caused it. And if a vendor does not know what a BAA is, you may already have your answer. ### 4. Ask What They Have Been Through Not what your dev team knows. What they have been through. Ask: ``` 1. Has any project of yours been through a SOC 2 or an ISO audit? 2. What came back? 3. What did you have to fix? 4. What has changed since then? ``` A team that has done it answers straight away, and do not worry if the answer includes something embarrassing. A team that has not will tell you about its security practices. Those are very different answers. ## The Takeaway None of this is hard. It is just early. Everything you cannot catch up on has to be running before you need it. Everything else can wait until Monday. Many teams do it the other way around: the visible half first, because it feels like progress, and the other half in a room with an auditor in it. ### Start-up Software Development: 3 Traps and 4 Things You Control URL: https://brival.co/insights/2026/08/startup-software-development Category: Methodology · By Mat Kupczyk · 2026-08-21 Building a first product rarely fails on engineering. It fails in the gap where a founder cannot tell whether the engineering is going well. Three traps account for most of it: building everything because it is cheap, building for a future you cannot see, and building for everyone. Here is what you control instead. Nobody warns a founder that running the build is a separate skill from having the idea. You hire people who seem good, you pay on time, and a year later you are looking at a product that does more than it should and less than you need. The video above is the full walkthrough. This is the written version, with the self-check questions in a form you can copy. ## It Is a Business Problem First The business is hard by itself, with no technology involved. Even when the product is the technology, a mobile app, an AI agent, whatever it does, it is primarily a business, and the struggles you hit are organizational rather than purely technical. That matters because it decides who has to solve them. A better framework will not make a decision nobody made. A faster team will not close a scope nobody closed. ## Why You Cannot Judge It Because you are wearing every hat, and nobody is an expert in all of them. Sales. Marketing. Onboarding. Then delivery, making sure the work actually gets done. Technology sits somewhere in that list, and it is usually the one you have the least equipment to judge. The closest comparison is a sales process you have never run. You do not know how to set it up. You cannot tell whether it is being handled well. Someone is doing it, but you have no tools to assess their craft, so you judge the outcome and nothing else. That is where most founders sit with their product, and they sit there against a clock. A start-up begins with a budget and a goal attached to it, not an infinite runway to validate. Every wrong thing you build comes at a cost, and if it does not earn its dollars it shortens your timeline.You do not need to know how the sausage is made. You need to know the frame you are operating in.
## Three Traps, Three Costs Each trap bills a different account, which is why the fix is different in each case. | Trap | What it looks like | Where the cost lands | | -------------------------------------------- | --------------------------------------------------- | ----------------------------------------------------------- | | Everything is cheap, so you build everything | Scope grows by small, reasonable-sounding additions | **You.** You can no longer explain your own product | | You build for a future you cannot see | Architecture designed against imagined requirements | **Your codebase.** A ceiling, then refactoring or a rebuild | | You try to serve everyone | Endless configurability, no depth anywhere | **Your market.** Nothing specific to choose you for | ## Trap One: You Build Everything The cost of building collapsed, and it took with it the thing that used to make your decisions for you. Development is no longer a matter of months and you no longer have to choose on price, which sounds like a gift until you watch it play out. Here is how it actually goes. In a new CRM, I need a list of people today. It is cheap to make it slightly more generic, so let us also add companies. Companies are in the system now, so an org chart is reasonable, for those enterprise clients, and it is low-hanging fruit. Then more ideas arrive. Their clients’ birthdays. Favorite colors. Their pets. All of it under the impression of justified business decisions, focused on the best possible experience. In reality they are distractions. You needed a list. Every one of those steps is defensible on its own, which is what makes this trap invisible from the inside. The budget was never a constraint you liked, but it decided what got in and what stayed out. Remove the price factor and every answer defaults to yes. So the cost stops being money. It becomes the complexity introduced into the product and the load of holding that complexity in your head. This is the thing I hear most often from founders working with AI: they are shipping so much, so fast, that they are no longer certain how their own product works. It lands on you, and on your ability to answer a client’s question about the thing you sell. We once took over a home-services app that had been live for a year, packed with features, and getting almost no real customer calls. The fix was subtraction. We [stripped it back to a single Call Now button](/case-studies/plunjr-mvp) to isolate the one question that actually mattered, which was whether customers would call at all. That answer had been buried under a year of reasonable-sounding additions. **Before the next feature, ask three questions:** 1. Did a client ask for this? 2. Did you schedule it upfront? 3. Or did it just look cool as you went? If it is the third one, you can probably table it. ## Trap Two: Guessing the Future Things are going to be unpredictable, and that is not a planning failure. You make an assumption, you act on it, then you pivot because the market moves you somewhere else. The problem is that building properly for the long term needs an idea of the future, and early on you do not have one that reaches far enough. Some products are exempt. Everyone knows a CRM has contacts, companies, opportunities and proposals. Those concepts existed long before you arrived, so assuming is safe. But if you are building something that has to be a CRM today, might become a warehouse system next year, and a gamification platform for HR the year after, you cannot make those assumptions upfront just to be prepared. There are too many unknowns. You may hit a ceiling in the codebase and need refactoring, or a rebuild of that part, to match what you have since learned. That is normal, and it is worth saying plainly: hitting the ceiling is not evidence that somebody did a bad job. It is the price of having learned something. What it is not is free. Unlike the first trap, this one is paid in code you already bought. **The practical version:** focus on the use case in front of you today. Ship what matters most to validate the next hypothesis, put it in front of clients, and adjust the plan from their feedback. That is how the future gets found. Do not fear the change, account for it. ## Trap Three: Building for Everyone This is the one I see most often, and engineers fall into it hardest. The instinct is to keep the product infinitely flexible, to handle every use case imaginable, while never diving deep enough into any single one. What makes a product worth buying is narrower than that. You understand a specific audience, you solve their specific pain, and you price it so that you make money while the client is glad to pay to have the problem gone. A product that is for everyone is really for no one. It is too generic, it does not stand out, and it is often so configurable that setup becomes its own burden. Technical sophistication, product excellence, total customizability: none of these introduce real value to your customer. They can make you feel better about what you built. They are not what the client is paying for. So do not spend on features, and do not lose sleep over them, unless you have a clear plan to turn them into revenue soon. This is exactly the moment a good vendor should tell you to stop building. Most will not, because building is the thing they bill for. ## What You Control You do not have to know everything, and you do not have to do any of it yourself. You have to understand enough to actively monitor the process. ### 1. Choose a Vendor, Not Hands Look for people who follow real practices, work efficiently with AI, and hold a quality standard they can point at. Whether it is an agency, a freelancer or a technical co-founder, pay attention to who they actually are, what they represent, and whether they have a track record with clients like you. ### 2. Keep Them in Check Use the right metrics and a reporting cadence. And now you can use AI, which is the part that genuinely changed: a non-technical founder can get a second opinion on their own codebase. Connect an agent to your code repository and ask it, in order: ``` 1. Have we set the right standard in place? 2. How was last week? What did we actually deliver, meaning what got baked into the code, not the progress presented in the report. 3. Here is the vendor’s report. Here is the actual code. Is it aligned, or is anything missing? ``` That does not make you technical. It gives you a second opinion you could not otherwise buy at that price. ### 3. Own What You Paid For Make sure what you are paying for is actually owned by you. You do not want the situation where someone hands you the work, you pay for it, and then you cannot sell the company because of a licensing obligation in a dependency nobody reviewed. We covered ownership and licences properly in [what a code audit actually covers](/insights/2026/08/code-audit-what-it-covers). The only thing to add here is: do not skip it. ### 4. Validate Before You Build Business is an unknown, especially early. With a brand new product you do not yet know whether it will resonate, which is natural rather than a failure of planning. So run discovery, validate the hypothesis, and put the thing in front of real clients to find out whether they can use it. Analytics on actual behavior helps. Do it fast and do it cheaply, because of the clock. One useful distinction: if you already have a validated service and you are adding software to it, you can afford to go wider, because you already know the shape of the thing. ## What to Expect From a Vendor This is what we learned works best for clients, so it is how we work. It is also reasonable to expect it from any vendor you hire. **They follow you, not their own habits.** Some clients are technical and want the engineering backlog; some just want the product delivered. Expect to be asked, or to start the conversation yourself, about how you stay in the loop and on which channel. **Frequent demos, not status reports.** Live, screen-shared demos are what keeps the feedback loop honest. Even with real attention on the agreed spec, there has to be buffer for necessary changes, because things always change. **A partner who optimizes for your business, not their billable hours.** You started with a budget and a goal, which is exactly why fixed price fits early stages. On [a recent academic CV product](/case-studies/ctrlcv-academic-cv-mvp), two non-technical founders took a clickable prototype to a production MVP in four months on a fixed price, with the scope set at the start and nothing renegotiated mid-flight. A good partner also raises distractions with you: when the work drifts from the goal, they say so out loud. **Support after the release.** The release ends the project for the vendor. For the business it is where things start. A small maintenance package, eight hours a month is enough to begin, keeps the system operational and buys you someone to answer questions, pull numbers, run a technical assessment and estimate the new work that real users will generate. **Changes treated as normal, not as an insult.** Pivots are part of the game, not an engineer’s nightmare. A good vendor adjusts to what the business needs instead of defending what is already built. But be honest about the mechanics: a material change mid-development, in a fixed-price model, means either extending the budget or pushing something else out of this phase. It is one of the two. **The right tool for the job.** Early on, do not over-invest in infinite scalability or expensive infrastructure. If you are building in healthcare or another regulated industry, do not cut corners on security either. This is why we put several approaches in front of a client with their trade-offs, and what each does to the goal, the timeline and the budget. Knowing the alternatives is leverage. ## The Takeaway You do not need to know how the sausage is made. You need to know the frame you are operating in, and to choose a vendor who understands your situation. Name what you are building and who it is for. Keep it small enough that you can still explain it. And put someone next to you who will tell you when you are drifting. ### What a Code Audit Actually Covers URL: https://brival.co/insights/2026/08/code-audit-what-it-covers Category: Methodology · By Mat Kupczyk · 2026-08-11 A code audit answers one question: whether your system can carry the weight you are about to put on it - ten times the traffic, an enterprise security review, a due diligence. It reads three things: how the team works, what you actually own, and how the system behaves under load. Not whether the code is “good”. An audit is not a grade. It is a forecast: an outside read of whether your system can carry a specific thing you are about to ask of it, delivered before you commit money to the answer. The video above is the full walkthrough. This is the written version, with the self-check list at the end in a form you can copy. ## “Is my code good?” has no answer There is no score. Every codebase has problems, including the ones we are proud of, so any verdict on general quality is an opinion wearing a report cover. The question that does have an answer names a load. Ten times the traffic. An enterprise client’s security review. A due diligence. Two years of a roadmap you already sold to investors. Once the load is named, “will it hold” becomes a technical question with evidence behind it, and the same codebase can pass one audit and fail another without a line of it changing. That is not hypothetical. A [healthcare platform we work with](/case-studies/healthcare-platform-modernization) ran well for years, then started winning enterprise clients, and enterprise clients bring their own security reviews. The code was the same on both sides of that line. It cuts the other way too. A consumer company came to us with a mobile app and an investor-promised B2B product they wanted built on top of it. They bought an audit first. It came back with serious flags on quality, maintainability and security, so the board postponed the product it had already promised, fixed the foundation, and shipped the new line on top of it afterwards. The audit never said the code was bad. It said the code could not carry that particular product yet, which is the only reason the product exists today. ([The full engagement](/case-studies/personal-safety-app-rebuild) is a case study.)Before commissioning an audit from anyone, write the sentence: in twelve months, this system has to ___. An audit without that sentence is a general opinion with an invoice attached.
## When an audit is worth buying Three situations account for most of the audits we run, and each has a load hiding inside it. **Something already hurts.** Releases are slow, things break, small changes take three times longer than they used to. Worth knowing before you start: what hurts is rarely where the problem lives. A slow checkout is usually not the checkout, it is the database underneath it or the way changes get tested before they ship. **You cannot grade the answers you are getting.** You ask how long something will take, or whether the platform is safe, and the answer arrives with no way to check it. This is the most common reason and the one people are least comfortable saying out loud. It is not distrust. You cannot grade an answer you could not have produced yourself. **You have committed to something you are not certain of.** The enterprise deal, the next scale milestone, the second product. The load is already named. Nothing has been measured against it, and because nothing is on fire, the cheapest moment to check is the one most companies skip. Sometimes the audit is not your idea at all: an investor before a round, a due diligence, a certification body, or you are the buyer and cannot read the code you are about to pay for. Below roughly the mid-market line the large diligence firms are not interested, so buyers at that size tend to guess. There are also three situations where we would tell you not to bother. If you are still searching for product-market fit, there is no load yet. If you already know the answer and want a document that agrees with you, the report will disappoint somebody. And if we cannot get access to the code, the tooling and the people, there is nothing to audit. ## Lens one: can this team carry the load? The audit starts with the process, not the code, because the codebase describes the past twelve months and the process determines the next twelve. Code can be replaced in weeks. The way a team decides, specifies, hands out and reviews work takes a year to change, which is why fixing a codebase without fixing the process just grows the same code back. A code-quality problem is very often a process problem wearing a costume. The most expensive thing in this lens is also the least asked about: where the knowledge lives. Not who is talented, but what stops entirely if one person leaves. In one due diligence, the entire product was running on the output of a single developer with no written contract. That is three findings, not one. Operationally, the company stops if he does. Legally, there are no terms and no notice period in either direction. And on ownership, the intellectual property had never been transferred, so the business did not own the product it was selling. None of it is visible in a code review, and in a diligence it is not an appendix item, it is the price. ## Lens two: is it yours, and can it take the next shape? This is the lens most people picture: the codebase and infrastructure read as an asset you own. **Architecture** determines whether your next change is priced in days or months, and whether the system survives a traffic multiple, a second platform, or the integration you just sold. **The data layer** is the expensive one, and the one nobody asks about. Code is cheap to change; three years of live records are not. If you hold personal data, this is also the first thing an enterprise security review reads. **Coding standards** decide how long a new developer takes to become useful, and whether an AI agent can safely touch anything at all. **Test coverage** decides whether you can change the system without gambling. Without a safety net, teams quietly stop touching the risky parts, and the risky parts are usually the valuable ones. That healthcare platform had effectively no automated tests on a system where a silent bug carries medical weight, and nobody had noticed, because nothing had broken yet. **Security posture**, in the code and the infrastructure, at the level that decides whether a deal stalls in someone’s vendor review. A dedicated penetration test is a different product, and we say so when that is what you need. **AI readiness**, which is two separate questions. Whether the codebase can be worked on by AI (standards, tests, documentation, structure), which sets your team’s speed for the next few years. And whether it was written by AI, quickly, with a foundation that never caught up. Then the two nobody asks for. **Licences**: every project stands on other people’s code, some of it carrying terms that expect you to publish your own source in return. A package three levels deep can leave you choosing between publishing your source and rebuilding the component. **Ownership**: paying for code is not owning it. Contractors, agencies, an early friend who helped for a few weekends, all of it needs a written assignment, or the work may not belong to the company at all. Both surface at the worst possible moment, which is a diligence with a buyer in the room. ## Lens three: what happens when the load lands The same system read as behaviour rather than as an artifact. **Infrastructure** tends to fall over on the day the product finally works. The quieter version is end-of-life: every framework and database you build on has a date after which nobody publishes security fixes. You did not choose that date, and it is already in the calendar. **Performance** is worth measuring rather than sensing. You paid to acquire the users who leave before the product gets good, so slowness means paying for the same traffic twice, while the hosting bill outgrows revenue. **Release** answers the most revealing question in the audit: can you deploy to production today? If a fix has to wait for a ritual, the company cannot respond at the speed its market assumes. The healthcare platform again: an end-of-life stack, no automated pipeline, releases done by hand, and no one aware of it because nothing had forced a test. **Monitoring** decides who finds out first. If you are not alerted, your customers are, and the worst case is not downtime, it is the site staying up while orders quietly fail. ## What you actually receive Every finding has four parts, in this order. | Part | What it means | | -------------- | ----------------------------- | | Finding | What is true about the system | | Evidence | The artifact that proves it | | Recommendation | What to do instead | | Cost to act | What that will take | The discipline matters more than the format. A finding without its artifact is an assertion, and a recommendation without a cost is not advice: “improve your test coverage” is the software equivalent of “eat healthier”, true and useless in the same breath. The report is written to be forwarded, because it usually is. It has to hold up in front of an engineering lead who will argue with the numbers and an owner deciding in money and months, and survive being sent to an investor without a translator.You are buying the order, not the list. Plenty of tools will hand you everything wrong with a codebase for free, and on that list everything looks urgent. The value is knowing what is first, what is next, and what you can live with for another two years.
The engagement itself is [discovery, analysis, a written report and a readout call](/services/code-audit). Discovery is where the load gets named. Analysis covers the code, infrastructure and tooling, plus conversations with the team, which is where lens one comes from and is not a formality. ## What to tell your developers Not behind their backs. It always surfaces, and the finding they will remember is that you went around them. Say what it is: not a performance review, a second pair of eyes on a system the business now depends on. The honest arithmetic is worth saying out loud to them, because it is genuinely in their favour. An audit that finds nothing is free reassurance for everyone. An audit that finds something hands them the case for the maintenance work they have probably been requesting for a year. Occasionally the real question is whether the team can do what is coming. That is fair, and an audit can answer it: we have written reports saying the team as it stands cannot carry what is next. What an audit cannot do is start from the answer. ## Five checks you can run yourself Here is the part that costs me something to write, as somebody who sells these audits: a good share of what we do in the first days, you can now do yourself with AI. 1. **Write the sentence.** “In twelve months, this system has to \_\_\_.” Until it exists, nobody can tell you whether your software is in trouble, including us. 2. **Ask for artifacts, not opinions.** Where are the coding standards written down? The test cases, the release runbook, the architecture diagram? You are not grading the answers. Absence is the finding. 3. **Ask what stops.** If this person left tomorrow, what halts entirely, as opposed to slowing down? 4. **Run the ownership check.** Does a signed agreement exist for everyone who has ever written code for you, including the early people? And can someone produce the licence list for every package you depend on? That list is an SBOM, a software bill of materials, and every package manager generates one in minutes. Most teams have never been asked for it. 5. **Time the boring things.** How long to onboard a developer, or a client? How long from a change being finished to it being live? That last one is lead time for changes, one of the four DORA metrics, and it says more about a company than a code review will. AI does most of the mechanical half. Point it at the repository and ask it to produce the architecture diagram nobody ever wrote, which tells you something either way. Ask which critical paths have no test behind them. Ask for every dependency with its end-of-life date, which is usually the most alarming list a founder reads that month. ## What the list cannot do for you Three limits, and they are the reason outside eyes still earn their fee. You cannot grade an answer you could not have produced. Your team cannot report what they do not know. And you have nothing to compare against: you have only seen how your own team works, so whatever it does is normal by definition. Every team believes its own workarounds are simply how software gets built everywhere. Which is the whole case for doing this before it is urgent. You do not audit to find out whether your code is good. You audit so that the next expensive decision stops being a guess. ### How We Made Design QA Measurable URL: https://brival.co/insights/2026/08/design-qa-measurable Category: Methodology · By Karol Kempa · 2026-08-10 Unit tests prove the logic. End-to-end tests prove the flow. Neither notices that a column header renders in the wrong weight, or that a row grew two points taller. This is the layer we added so something would, and what it changed. ## What our other tests were never looking at Unit tests prove the logic is right. End-to-end and integration tests prove a flow completes. Neither of them notices that a row grew two points taller, that a separator changed shade, or that the label on a button appearing across five screens changed size. That is not a failing of either. They are not looking at the rendering: they ask whether the app did the right thing, and they answer that well. Checking the rendering by hand is possible, and it is not even hard. An inspector will give you a font size to the point: open the screen, find the element, read the value. You just have to do it for every element, on every screen. And by the fifth pass over the same screen, change blindness takes exactly the details you came for. So we added an instrument that answers two different questions. The first: does this value in the code come from the design system. The second: has the rendering changed since last time. Neither replaces the other, and neither tells you whether a screen looks good. The two questions, and the two halves that answer them. Everything on the left asks whether a value belongs to the system; everything on the right asks whether the rendering moved.
## The flow, end to end It starts in the design file, and that is the condition everything else rests on. Color and type have to be variables rather than values pasted onto a layer. Only then does a value have a name that can be read mechanically. Screens have to be numbered frames; we have fifty-eight of them across thirty-six names, so every screen has an address you can ask about. Our designers built the file that way before anyone thought about linting anything. Figma exposes every frame over an API. For each element you can read its typeface, size, weight, letter spacing, color and position: the same properties the code sets, in the same units. The fix was to stop treating the design as pictures and start treating it as data. And because every screen has an address, we can check whether it matches the design. Pull that frame's nodes and compare element by element against what the code sets. A screen takes a few minutes and the output is a table rather than an impression: this element, this is what it renders, this is what the design says. One screen read out of the design file and compared element by element against what the code sets. Four deviations, none of them catchable by eye.
The variables come down into the repository as a token set: twenty-three colors and eighteen type styles. The result is committed, so the history of that one file becomes the design-system changelog nobody otherwise writes, and the reference lives in the repo: nothing has to reach the design tool during a build. A static scan reads the app against that token set. Interface files, source, color catalogs. It reports every color off the palette, every size off the scale, every spacing off the grid and every letter spacing that disagrees with its token, each with a file and a line.The scan's output: one row per rule, with severity and count. It needs no simulator and finishes in seconds.
Next to it sits a visual layer. Screens and components are built in a test off mocked data, rendered to an image, and compared against an image committed in the repository. Today that is a hundred such images (forty-one of screens and fifty-nine of components) covering nineteen screens and twenty-two components, because one screen usually owns several images: an empty state, an error state, a variant. The comparison runs at three moments. Locally, in one command, before you open a pull request. On the pull request, where a single run first compares the images and then builds the report, comments it, and writes a summary into the description. And between releases: because the reference images are committed, "the previous version" is simply those files at a tag, so comparing two releases is not a build at all: nothing is rendered, no simulator starts, both sides already exist. The two layers answer different questions, which is why neither substitutes for the other. The scan will tell you a color is not in the palette, but not that a screen looks different from last week. The rendering will tell you something moved, but not whether the value that moved exists in the system at all. ## What makes the comparison mean anything The data is mocked, so both sides get the same input. Every pull request therefore starts from the same structure, and two runs a week apart compare cleanly. The rest of the repeatability is structural rather than a matter of discipline. Doubles in place of services, pinned image geometry, a pinned time zone and locale, guarded by a test of their own. There are more component tests than screen tests, deliberately. Sensitivity is inversely proportional to the size of the thing you compare. The same fix (the one in the table above) moved under one percent of the whole-screen image; measured against the cells it actually lives in, one and a half to three. A change to a control that appears on five screens dissolves into each of them, so the things we care about are asserted as components rather than as the screen they sit on. What the report looks like at screen level: reference, current render, and the pixels that differ, painted red.
The same report at component level. The smaller the thing asserted, the larger the fraction a given change occupies.
## What changed in the work A change surfaces before review, not after. An engineer runs the comparison locally in one command and sees which images they moved before anyone starts reading the code. They fix it in the same sitting, instead of coming back to it two days later on someone else's comment. The review question became answerable. Every pull request carries a comment naming what changed, linking to before-and-after images the platform renders as a slider, and a block in the description naming the screens behind them. The question stopped being "did this touch the UI?" and became "this changed four screens; are those four intended?" Engineers and QA see the same evidence at the same moment, before anything is installed anywhere. Accidental changes stopped being invisible. This is the payoff we went in for. A change to a shared component spreads to screens its author never opened and had no reason to open, and the report names them, including the changed screens that no test renders at all, because knowing where you have no evidence is worth more than another green check. QA stopped hunting and started judging. Nobody sweeps screens looking for two points of difference any more; they get the list of what moved and decide whether the move was intended. The machine does not get bored on the fifth pass over the same screen. A release gets one page instead of a conversation afterwards. A release is an accumulation of merges, each reviewed on its own. Comparing two tags takes seconds and tells you what to talk about before shipping. ## Where it stops The scan tells you whether a value comes from the system, not what it should be. The design says that, which is why both halves are needed. The rendering tells you something changed, never that it is wrong. It covers nineteen of the app's thirty-one screens, which is why the report names the ones it does not: knowing where the evidence is missing is part of the result rather than a hole in it. Deciding whether a change was the one you meant to make is still a person's job, and we do not expect that to move. What changed is when and on what that decision happens: on one page with both images side by side, before the release, rather than in a message from a client three weeks after it. ### How AI Changes Every Step of Building Software URL: https://brival.co/insights/2026/08/building-software-with-ai Category: Methodology · By Mat Kupczyk · 2026-08-04 AI didn’t change how software gets built - it changed what each step costs. Across all seven steps, from strategy to monitoring, work that used to need a budget line now takes hours: research, prototypes, tests, pipelines, reporting. If your team pays for AI tools and you feel no difference, the gap is in the steps, not the tools. If you are wondering whether AI is enough of a support to feel a difference in how your dev team performs, my answer is that it is here, and you should notice it by a mile. If you don’t, something specific is off - and it is findable. Software still gets built in the same seven steps it always was; what changed is what each of those steps costs. The video above is the full walkthrough; this is the written version, for skimming and for sending to whoever runs your product. ## Your team pays for AI tools. Why don’t you feel a difference? Because the tool alone isn’t the change. AI speeds up the steps of building software, so if a step isn’t really run - precisely and deliberately - there is nothing concrete to speed up. It is a multiplier, and a multiplier works on whatever you feed it. We see [software as built in seven steps](/insights/2026/07/software-delivery-flow), and those steps haven’t changed. Two things about them are new. First, every one of them should be dramatically enhanced by AI today. Second, each is ultimately meant to be taken over by agents later. Here is what changed in each, and one practical thing your team can start doing tomorrow. ## 1. Strategy: research that used to take days now takes an afternoon Strategy answers why we are building something, before anyone builds it. The old way: decisions got made on a hunch or on vague signals - the loudest customer, or the competitor’s last launch. Meaningful research meant days of work or a consultant’s invoice, so mostly it just didn’t happen. Now the knowledge is on tap. You can have AI research recent trends, the competitive landscape, and what similar products charge; for less than a hundred dollars you get a comprehensive report that would have taken days to prepare. We run these for our own decisions now, not just for clients. It works on the behavior side too - before you dive into the future, you should understand the current state. Open your codebase and ask how to pull numbers on user behavior: what people use and don’t, how they click through the app. That is how you stop guessing what your product does for people and start knowing.Before the next big feature, do both in one sitting. One research report on the market, one answer on how you’ll measure usage. Half a day, total.
## 2. Requirements: a clickable prototype instead of a wireframe wait Requirements is where the idea becomes specific enough to build and to verify. We are building a mobile app where someone can sign up, log in, see a list of products, and order one - that easy. The old way included prototypes, designs, and week-long design sprints. Meetings, and then more meetings to make sure everyone understood the same thing, because building the wrong thing costs thousands in budget and months in time. And then you waited days for a wireframe. The new way: capturing requirements doesn’t have to be a post-session effort, because you can collect them as you go. Getting clear on the shape and the visual has never been easier, and nobody waits days for a wireframe or a spec.Invite an AI assistant to the product meeting. Take the transcript, upload it to Claude, and have it aggregate the notes. Then put those notes into Claude Design - a tool that turns written requirements into a real, clickable design - and describe what you want. You get a clickable prototype within hours. A non-technical, product-aware person can build it, click through it, and correct it in chat, with no design skills needed. When it is right, it goes to the developers as the starting point for the build, converted to code their AI can pick up.
## 3. Architecture: the options that used to need an enterprise budget Architecture is the structure underneath the product - how the pieces fit together, and what decides whether the next change is cheap or painful. The old way: changes here, or any of the more careful patterns, were probably too expensive. Structure got chosen by habit, there was no time for refactoring, and the documentation never got written, because all of it was overhead nobody budgeted for. Now that overhead is basically free. Diagrams and documentation take minutes and live with the code. Approaches that used to need an enterprise budget are on the table for a regular team, so you can think through options that used to be out of reach: an MVP that integrates with more than one provider, handles real traffic from day one, launches on more than one platform.You don’t need to know a hundred design patterns by heart, you need to ask the right question. Open the codebase and tell AI: we are building X, we see it potentially going in the direction of Y and Z, can we structure it in a way that supports X now and keeps things open for Y and Z? Then let your senior person judge the trade-offs. The knowledge is free now; the judgment still comes from experience.
## 4. Coding: quality stopped being a budget question Coding is writing the software to one shared standard. The old way: coding was the most hours, the largest team, and the biggest bottleneck. Everything waited on hands and keyboards, and when budgets got tight, quality was the thing to cut - few tests, no documentation. For us the new way is the obvious one, because AI writes all of our code now, and does it better than us. But the real shift is what that does to quality. Tests and documentation stopped being a budget question, and the big cleanups that used to be [a quarter of a roadmap](/insights/2026/06/what-ai-changes-in-a-software-estimate) happen at a fraction of the old effort. I have seen a single strong engineer, with AI and a clear process, outperform a team of three or four - at a quality most companies couldn’t afford two years ago.Write your standards down inside the repository - naming, structure, how much testing ships with each feature. The AI tools read those files and follow them. That is the difference between AI as a faster typist and AI as a consistent team member.
One honest caveat: someone on your team must still understand the code. A codebase nobody holds in their head, at least on some level, is a terrible place to debug. AI writes, but your engineers stay accountable. ## 5. QA: test automation is close to free Quality assurance is about catching problems before your customers do. The old way: test cases and automations existed, but not many teams had the budget to keep them extensive and up to date, so it came down to manual effort. The developer finished, and a tester clicked through the site to confirm nothing broke. Or nobody did, and your customers were the first ones to find out. The new way is simple. Test cases and automations are close to free, so there is [no longer a reason to trade one against the other](/insights/2026/07/qa-testing-documentation-trade). Testing the app should be documented and, preferably, not fully performed by a human.Ask AI to write test automations for your critical paths - sign-up, login, order, payment. Then every time you finish a change, run them and let the robot click through the site to see whether anything broke.
## 6. Release: releases should be boring Release is getting the change live safely, predictably, with everyone who needs to know informed. The old way: in most projects, releases were events. Scheduled after hours, with a manual list of steps if you were fortunate, one person who knew how to run it, and everyone holding their breath until Monday. The new way enables quick iteration. The automation pipeline that used to need a dedicated specialist now gets built in an afternoon. Releases should be small and frequent - even a few times a day if you want - and the release documentation writes itself: what shipped, when, and why. Releases should be boring, and that is the goal.Ask AI to build the pipeline and a short smoke-test checklist - the five things that must work after every deploy - and let the automation do the checking.
## 7. Monitoring: the daily report nobody would have paid for Monitoring is knowing what is actually happening once the software is live: whether the system is up, and whether the business is working. The old way was often flying blind. Dashboards were an infrastructure project, so early products skipped them, and you learned about problems from customer emails. Now the cost of integrations and monitoring is marginal. Have your team expose an endpoint with the application status - signups today, orders today, errors today - and wire those numbers into daily reporting. You probably wouldn’t have paid for eight hours of dev time to send you a daily report on Slack; now it takes about thirty minutes, published through whichever channel you want and customized to whoever reads it. And when something looks off, you can ask questions in plain language instead of digging through dashboards.One status endpoint, one daily message. That is the whole build.
## Seven questions to ask your dev team One question per step. Take them to your next conversation with your dev team - not as an audit, just out of curiosity. Each answer will tell you a lot about how well that step is handled. 1. **Strategy** - what’s the goal of the thing we’re building right now? 2. **Requirements** - where can I see what we’re building, before it’s built? 3. **Architecture** - which part of the system are we afraid to change? 4. **Coding** - where are our standards written down? 5. **QA** - where do we have our test cases? 6. **Release** - can we deploy to production today? 7. **Monitoring** - what are we alerted about when things fail? A good team will enjoy these questions. If some of the answers are missing, you just found the step to start with. It is the same standard we put on the buying side in [our vendor checklist](/insights/2026/07/software-vendor-checklist) - these are those questions asked from the inside instead. ## AI didn’t replace the discipline. It removed the excuse. Same seven steps we have always used. Research in hours, prototypes in a day, tests almost free, releases that are boring, and reports that come to you. That is also the direction this is going. Today AI enhances each step, but over time [agents will take over more and more of each one](/insights/2026/07/self-learning-apps). Teams that run a clear process will hand it over smoothly. Teams that run on instinct will have nothing to hand over. So send this to your dev team, or whoever runs your product, and ask them one question: which of the seven steps do we start with? ### Self-Learning Apps: What It Takes for Software to Improve Itself URL: https://brival.co/insights/2026/07/self-learning-apps Category: Methodology · By Mat Kupczyk · 2026-07-31 Software that improves itself is buildable sooner than most founders think, but it is not a model you train. It is three levels of readiness: a system agents can operate safely, processes written down instead of living in people’s heads, and a codebase disciplined enough to let an agent ship changes. A few weeks ago, a client of mine, a start-up CEO from the US, asked me how to make his system self-learning: software that improves on its own, so the competition never catches up. He is not alone. The founders he talks to want the same thing, and honestly, I find the idea inspiring. But between the dream and what your software can do today there is a gap, and it sits on three levels: the system, the process, and the codebase. The video above is the full walkthrough; this is the written version, for skimming and for sending to whoever owns your roadmap. ## Level one: the system has to give agents a way in For AI to work inside your product, it needs a communication layer built for agents. Most systems don’t have one, and that, more than anything else, is the bottleneck. The simplest way to see it: ask Claude to send an email, and it does it nicely, because Gmail plugs into AI. Now ask it to create a lead in your ten-year-old custom CRM, and it will tell you it just can’t. The difference is not intelligence. Gmail gives agents a way in; your CRM doesn’t. And it doesn’t matter whether the thing knocking is a bot, an automation, or an integration: they all need the same thing. The tempting shortcut is “we already have an API, just plug the AI into that.” Even Anthropic, the company behind Claude, [lists wrapping existing endpoints as a common error](https://www.anthropic.com/engineering/writing-tools-for-agents) in its guidance on building tools for agents. This layer is its own engineering. You design it differently, you test it differently, and you add things most systems simply don’t have today: guardrails and logging. You decide what the agent may do on its own, what needs a human click, and what it may never touch. And you log every decision, what was done and why, because when an agent refunds the wrong customer, “the AI did it” is not an incident report. So before AI can work in your product with any freedom, the product has to be ready for it: either rebuilt to the new standard, or with a new layer built on top of it. ## Level two: the process can’t live in people’s heads You can’t hand a process to an agent if the process lives in people’s heads. Writing it down is not bureaucracy; it is the delegation itself. Try this one yourself. Ask ChatGPT to write a proposal for a client from the accounts of two different people, and you get two fundamentally different documents: different length, different language, different content. But give it a documented process with a template, and it puts the effort in the right places. It preserves the spine, adjusts the copy, alters the price, and replaces the client’s details based on the call transcript. It doesn’t drift away from the rules you set, and that is the whole point. If the edge cases get handled because one person just knows what to do, there is nothing to hand over. My test: if you can’t delegate it to a smart intern with a one-page checklist, you can’t delegate it to an agent either. AI is a multiplier, and it multiplies whatever you feed it. Feed it a fuzzy process, and it will execute the gaps at machine speed. ## Level three: the codebase has to survive an agent’s changes A product that improves itself means an agent commits code to your product. That is only safe on a codebase with tests that work as guardrails, docs an agent can actually read, and releases that run on their own. Here is what the full dream looks like. Your customers keep sending one extra detail every time, right after they sign up; it shows up in every support thread. The AI notices the pattern and decides this information belongs in the sign-up form. It adds the field, ships it, and then checks: did those support messages stop? That is a product improving itself. Nobody wrote a ticket, nobody planned a sprint. The product saw its own gap and closed it. It really is that good, and to be fair, this layer is the youngest of the three. If somebody promises it to you next sprint, that is a tell. Because here is the catch: the agent will make that change either way, confidently, rating its own work ten out of ten. If your codebase has tests and docs, the change ships and the tests prove nothing else broke. If it doesn’t, nobody knows what just broke three screens away: not you, not the agent. Your customers find out first. You can’t delegate changes on a codebase that fights every release. What makes it safe is the same discipline that makes human teams fast; agents just make the gaps show up sooner. ## What’s real today, and one number to keep you honest The technology is here. We have used it, and it is capable of everything above the third level’s frontier. What I haven’t seen is many companies using it in a meaningful way yet, because it takes a different approach and a different skill set on all three levels. The number: [Gartner looked at this in June 2025](https://www.gartner.com/en/newsroom/press-releases/2025-06-25-gartner-predicts-over-40-percent-of-agentic-ai-projects-will-be-canceled-by-end-of-2027). Of the thousands of vendors calling themselves agentic, they estimate only about 130 actually are; they even coined a name for the rest, “agent washing.” So when you hear a big promise, walk it down the three levels: the system, the process, and the codebase. It is the same discipline we put on the buying side in [our vendor checklist](/insights/2026/07/software-vendor-checklist): the AI question and the handover question are this standard seen from the other chair. As for us, we are on our path to the Claude Partner Network certification, keeping up with the newest in the space, so we can help tech companies get up to speed on all three layers. If you want an honest read on where your product stands today, that is exactly what our audit is for. ### Test It or Document It: The QA Trade That AI Ended URL: https://brival.co/insights/2026/07/qa-testing-documentation-trade Category: Lessons · By Tomasz Majcherczyk · 2026-07-30 For years QA made the same quiet trade on every project: test the product, or document it - never enough time for both. Six months of AI ended that trade for our four-person team. On one production project, written test cases went from 100 to 460 in two months; the judgment at the center of testing stayed manual. ## The trade we used to make Before any of us had Claude in the loop, the work was tedious in a specific way. Testing itself was fine - that is the job. The problem was everything the job needs to stay healthy: test cases written down, documentation kept current, tickets that a developer can actually act on, a project in good enough shape that the next person isn't starting from zero. Some context first. There are always multiple projects running at once - some in maintenance, some in heavy development. Most are startup products, and startup products breathe: a product ships, then sits quiet for months while real users generate the feedback that shapes phase two. Delivery QA is four of us, and nobody owns just one project - when a product goes into its quiet phase, you move to one that's live, and when phase two kicks off, you come back to it. That rhythm is why the trade was so hard to avoid: your week is split across products at different stages, and every switch means loading a different product back into your head. There was never time for all of it. So every week you made the same choice: test the build in front of you, or keep the documentation honest. Testing won, every time - testing had a deadline, and the product in front of us always got tested. It was the writing-down that slipped. It lived on Google Drive, it went stale, and maintaining it was a chore nobody could afford. Multiply that by four testers and a dozen-plus live projects and you don't have a documentation problem, you have a documentation debt that compounds across the whole portfolio. That is the part outsiders miss about QA capacity. The bottleneck was never the willingness to document. It was that the supporting work and the testing work drew from the same small pool of hours, and the supporting work always got cut first. ## First, we used it like a chatbot The first phase was cautious and obvious. "Claude, tidy up this ticket description." "Claude, draft me a test case for this." It helped, the way a faster keyboard helps. Low stakes, small wins, nothing about the shape of the work changed. We were using a chatbot on the side. That phase had its own tax, though. Unguided, the model padded everything it touched - bloated ticket descriptions, test cases with hallucinated steps for screens that didn't exist. At first we used it anyway and stripped the unnecessary content back out, which quietly ate the time it saved. It took a while to learn what actually makes it faster: giving it the context and the standard up front, instead of editing its guesses after. ## Then we stopped chatting and started building The shift that mattered was giving up on chat. Instead of asking an assistant the same thing over and over, we built a set of skills - small, standard-aligned tools that hand work down a chain, so the tester supervises instead of retypes. - **The tc-generator skill** turns source code into human-readable test cases, written for a manual tester - what they click, what they see, what they verify. No selectors, no endpoints. - **A ticket-generator skill** turns a rough, Slack-style note into a clean, correctly formatted bug ticket. - **A project-review skill** runs our QA checklist across a project's tickets and documentation and hands back one consolidated report. The move from "we use Claude for descriptions" to "go Claude, faster" was exactly this: from asking a chatbot for text to running skills that do a QA task end to end against our own standard. Changes that used to sit on the backlog for weeks now happen in an afternoon, and at a higher quality than we managed by hand, because the standard is baked into the tool instead of living in someone's memory. ## The tool underneath it: TCMS Skills produce test cases fast, but scattered test cases are just a faster mess. So we built an internal tool - [TCMS](/tools/test-case-management?utm_source=insights&utm_medium=post&utm_campaign=qa-testing-documentation-trade) - that pulls test cases together across every project. Testers run them, and the tool keeps the history: when a case last ran, whether it passed, what changed. That's the part that turns generated documentation into documentation we actually keep. A test case there has a history behind it, not just a file that got written once and never looked at again. The numbers moved in a way that still surprises me. Before AI, thin budgets meant thin written coverage almost everywhere - a project would carry a handful of documented cases while the real coverage lived in the tester's head. That worked, until the rhythm of startup work tested it: a product pauses after release, the tester moves to a live project, and months later phase two starts with the coverage waiting to be rebuilt from memory. On one production project, zero to a hundred written test cases took years of regular development; a hundred to 460 took two months. Some of the generated cases are smaller in scope - but they're maintained, tracked, and organized in a way the hand-written hundred never were, because the skills write them to one standard and TCMS keeps their history. It's not that anyone suddenly found the discipline; it's that a first pass from the source now takes an afternoon, not a week. The floor came up everywhere at once, on every budget - and the project nobody had time to document is often the first one we document now. ## One standard, every project The last piece was consolidation. We rewrote and merged our QA SDLC document and the standards around it - test cases, documentation style, commits, the way of doing things - into one written source of truth. That sounds like housekeeping. It is the thing that lets AI help without drifting: when the standard is one document, a skill on Project A and a skill on Project B produce the same shape of work. Consistency across projects stopped being a matter of who happened to be testing. ## What our meetings turned into The clearest sign the change was real showed up in our team meetings. Four testers across a moving set of projects means we sit down together regularly, and for years those meetings ran on the same fuel: technical questions asked over and over. How do we name this. Where does that doc go. What tag do we use here. Which format for that ticket. The answers existed - in someone's head - so every meeting spent its energy re-deriving the way of doing things instead of moving it forward. Now the flow is written down and the tools carry it, so those questions mostly don't reach the room anymore - you ask the standard, or the skill already did. What's left is the meeting we never used to have time for: the bigger picture. We talk about direction, where we want QA to go, which new tools or approaches are worth a look, what's working and what isn't. In practice we gave a small QA department its own R&D loop - time to research and try things instead of just re-explaining them. That's the part I didn't see coming, and it might be the most valuable thing we got. ## What did not get easier None of this touched the actual test. AI drafts the test case, but a person still has to decide whether the case is the right one, and whether the product in front of them actually clears the bar. That judgment is still manual, still the slow step, and honestly it should be - [we've measured what that means for a project's budget](/insights/2026/06/what-ai-changes-in-a-software-estimate?utm_source=insights&utm_medium=post&utm_campaign=qa-testing-documentation-trade). The judgment isn't the only thing we haven't solved, either. A few pieces of this are still unfinished. Mobile is where the tooling runs out. Our browser-driven exploration doesn't reach a phone, and the agent that would is still an experiment - not something we'd point at a production build tomorrow. On mobile projects we're mostly back to doing it by hand. And we've defined the metrics that are meant to show QA's value and started feeding them into a weekly board - but we're at the start of that trend line, not the end. Some of it we can't even measure yet, not until the trackers carry the right fields. I'd rather say all that plainly than pretend the last six months tied a bow on it. ## What we actually gained The headline everyone reaches for is "AI makes the work faster." For QA, that undersells it and slightly misses the point. What AI removed was the trade. The documentation, the test cases, the tickets, the standards - the work that used to lose every time it competed with testing for hours - is now cheap enough that it stops competing. Quality across a project became something we maintain as a matter of course, not something we had to ration hours for. For a four-person team spread across a shifting set of client projects, that hits in a very concrete place. When someone picks a project back up after its between-phase pause, or takes one over from a teammate, the documentation is actually there and actually current - the handoff takes minutes, not days. The thing that used to make us fragile was knowledge living in one person's head on one project. That's exactly the thing that got cheap to write down. We didn't fix that by hiring; we fixed it by making the knowledge cheap to keep outside our heads. The judgment didn't shrink - it's the job, and no model does it for you. What changed is everything around it. And once you've worked that way for a while, you plan QA differently: you stop spending your energy defending the supporting work, and start spending it on the calls that were actually yours to make. ### How to Choose a Software Vendor in 2026: The Complete Checklist URL: https://brival.co/insights/2026/07/software-vendor-checklist Category: Methodology · By Mat Kupczyk · 2026-07-22 You will never judge a vendor’s code before you sign — but everything around the code, you can. Seven questions before signing, four things that must be in writing, three signs of a healthy engagement, and three rules for when the vendor leaves. It’s the standard we run our own engagements to. You’re about to trust a software vendor with your product — and you can’t judge their code. The good news is you don’t have to: everything _around_ the code can be judged before you sign, and that’s what this checklist covers. The video above is the full walkthrough; this is the written version, for skimming — and for sending to whoever signs. One note on words: we say “agency” throughout, but the checklist is the same whether you’re considering an agency, a freelancer, or an offshore team. It applies to anyone who builds software for you. ## Why founders get burned Founders don’t get burned because they’re careless. They get burned because they can’t judge code quality — and because the things that will actually hit them sit _around_ the code, where nobody points: the deployments, the ownership, the security, everything that happens after the handover. So if you’ve been burned before, it probably wasn’t your fault. You engaged an expert, you judged what you could see, and the important part was invisible. We’ve seen this from all angles — we’ve worked on projects coming from agencies, delegated work to agencies, and taken over from agencies. And honestly, some of this list we learned by getting it wrong ourselves. Our standard today is a list of things we decided never to let happen again. ## What changed in 2026 AI changed this game in two directions at once. On one side, it collapsed the old trade-off between good and fast — but in our experience, only for teams with real discipline. Test automation and solid architecture used to be a budget discussion; they’re now cheap. Which means you should be demanding **more for the same money — not the same for less**. On the other side, the same AI lets an undisciplined vendor produce bad software faster than ever. Steve Krouse put it best: _vibe code is legacy code_. You can now buy a large, buggy legacy codebase in record time. So AI means you’ll be getting more either way — more value, or more debt. ## When an agency is the wrong answer Sometimes it is. When the software is your core product, long-term, and nobody on your side will ever own the tech — Y Combinator has warned founders about that setup for years, and they have a point. And when your budget is under five thousand dollars, hire a skilled freelancer; the agency overhead won’t serve you. The rest of this checklist is for when an agency is the right call. One thing before the questions: none of them are rude. A good vendor has heard every single one and doesn’t mind, because they have answers. If a question makes the room awkward — that’s feedback too. Speaking as a vendor: the clients we’ve worked with the longest are the aware ones, the ones we’ve had those “uncomfortable” conversations with. ## The seven questions to ask before you sign 1. **They diagnose before they quote.** In the first meeting, give them room and watch what they ask about. The best ones start with your business and your goals, not the project. Be careful with anyone selling a rewrite before they’ve read your code — that’s selling, not diagnosing — and with anyone whose favorite stack is the answer to every question. The strongest good sign of all: they tell you when their own solution would _not_ work. 2. **Deliverables on paper.** Know exactly what you’re buying: an output (working software), an outcome (a business result), or hours. And what you’re paying for: the result, the effort, or availability. Each is a fair deal — but each carries different risk, and it has to be written down, not assumed. We’ve taken over projects where the budget was drained at fifty percent done, and that’s when the client learned they’d been paying for hours all along. 3. **How they use AI — and what it changed.** A vague “AI makes us ten times faster” deserves a follow-up: how, and what did it change? Some vendors don’t use it (falling behind); some use it quietly and keep the difference as profit; some pass the savings on and deliver the old thing, cheaper. Our take: even when the work takes the same time, it should be done _better_ — better quality, better scalability, better architecture. Real buyers are pushing on price too — KPMG pressed its own auditor into a double-digit discount by pointing at AI savings — but notice where discounts land: at firms billing by the hour. Decide what you’re buying; cheaper is on the table, but better is usually the bigger prize. One thing here that points forward: making software friendly to AI — easy for models to develop, and for agents to interact with — is becoming its own engineering discipline, the same kind of shift as moving to APIs ten years ago. If you expect to need it, make sure the codebase will be ready for it. Good answers in this section sound concrete: “we didn’t have budget for end-to-end tests — now we do,” or “we build things to be AI-friendly.” New capabilities — not the old stuff, cheaper. 4. **Who carries the risk.** With a fixed price, the vendor carries the delivery risk. On time and materials — and here’s the honest part — _you_ own the delivery. Nobody says that out loud, but that’s the deal you’re signing. The fastest way to tell the difference is the warranty question: who pays for the bug that shows up two months after launch? If the contract doesn’t answer it, you just did. 5. **The handover standard.** When the engagement ends, could a competent stranger — or an AI agent — pick up the codebase next month and safely make a change? Tests, docs, automated deployments: that’s what the handed-over package should include. Worth adding: “if we grow an in-house team, can you work alongside them?” Your local team is going to grow, especially on the funding journey. 6. **Their delivery standard, end to end.** What does QA look like? How do they handle security and keep the platform reviewed and updated? What does maintenance look like after launch? Here’s the trick: you’re not judging the content of the answers — you’re checking that rehearsed answers exist at all. 7. **Start small.** A fixed-price, closed-scope first engagement with milestone payments tied to deliverables you’ve accepted. It closes the loop at small scale — how you communicate, how you agree on expectations, how delivery actually feels. For team augmentation: a two-week trial ending in a clear deliverable. The trial tests everything the first six questions asked about — against reality, instead of answers. ## The paperwork: four things that must be in writing 1. **Paying is not the same as owning.** Only a written transfer of intellectual property makes the code yours — access to the repository is not legal ownership. Check what the vendor keeps: their pre-existing tools and components, which is fair. Protect your IP, and let them protect theirs. There’s a saying: one client in a vertical is a client, two is a conflict, three is specialization — know where you’re leveraging their expertise, and land on fair terms. 2. **Acceptance.** Agreed criteria and a defined window to review each milestone — usually around five working days — with payment releasing on written acceptance. That’s what makes milestones real. One caveat: when work is sequenced, dragged-out reviews disorganize everything, so negotiate a healthy buffer and expect pushback if your reviews drag. 3. **An exit clause.** Termination for convenience — roughly thirty days’ notice and a capped exit fee. It’s the clause that lets you leave without a lawsuit. 4. **Warranty, with numbers.** The market norm is thirty to ninety days after acceptance for real defects, sometimes up to one hundred eighty. That’s your answer to the bug two months after launch. ## Three signs of a healthy engagement You can’t always review the code — so watch the delivery instead. 1. **Working software, delivered early and often.** Weekly or bi-weekly demos of running software, not slide decks about progress. A big final reveal after months of work is a big final risk — yours. 2. **They build in the open.** Commits and pull requests landing where you own them, as they happen — not “it’s almost ready, we’ll commit later this week.” Visibility isn’t a report; it’s watching the work land. 3. **Communication keeps you in the loop — and in control.** A named delivery manager, product demos every sprint, budget and scope check-ins. Scope changes surfaced ahead, with their cost and date impact — not at the invoice. And the standing ability to change direction, to the extent you agreed. ## Three rules for when they leave Every engagement ends, and the good ones are designed to. 1. **Ownership transfers on a milestone — not “at the end.”** A vendor holding ownership until they’re paid is fair; tie the transfer to the first paid milestone, not to the end of the relationship. From that point: repository, cloud, domains, pipelines — in your name, with the vendor holding access you can revoke. Clients do get held hostage — by code, by infrastructure, by a domain. We’ve seen a client strong-armed into paying extra just to get their own domain back when paths split. If nobody can state the ownership terms, you’re locked in by default. 2. **Make the knowledge transfer a deliverable.** Handover docs and recorded walkthroughs, written into the scope — because when the contract ends, the understanding walks out the door unless you caught it first. 3. **Leave-by-design is the green flag.** A vendor whose business survives you leaving has their incentives aligned with yours. Lock-in is a business model — not an accident. ## The pattern Everything on this list is one idea seen from different sides: **discipline you can observe, risk on the vendor’s side, and ownership on yours.** When those three hold, the details take care of themselves. And the reason it matters right now: every few years, software has to be re-done, and the shift this time is that systems are becoming agentic — built for AI to operate them, and open enough to plug into everything around them. The vendor bar moves with every wave. That’s why the handover question — and the quality question — belongs in every vendor conversation this year. ### Stabilize, Rebuild, or Start Over? How to Decide What to Do With a Legacy Codebase URL: https://brival.co/insights/2026/07/modernize-legacy-codebase Category: Methodology · By Mat Kupczyk · 2026-07-15 When a codebase starts holding the business back, you have more than two options. There are four: continue as is, stabilize, strangle incrementally, or a full rebuild. Which one is right has less to do with how bad the code is, and more to do with whether you can change the system while it keeps running. So measure first, then choose. Sooner or later, the codebase that got you here starts holding you back. Changes take weeks. The features customers keep asking for feel out of reach. Or a security review blocks a deal. Then you have a call to make: patch it, rebuild it piece by piece, or tear it down and start over. Most teams make that call on a gut feeling, and it's one of the most expensive coin-flips in the business. The video above is the full walkthrough. This is the short written version, for skimming and reference. ## The five symptoms that say it's time Modernization rarely starts as a plan. It starts as symptoms. The five we run into most often: - **Scalability and performance.** The system hits a ceiling. Sales keeps closing clients it physically can't carry, so you either turn away revenue or crash under the load. - **Operational struggle.** It works, but it's closed off. No clean integrations, no interfaces. Every enterprise you onboard costs a week of engineering just to wire it up to anything. - **Poor architecture and too much complexity.** One change breaks five other things, integrations get unreliable, and developers start routing around the system with workarounds. Often the fix is removing complexity: folding ten fragile third-party integrations into code you own, or collapsing three separate front-ends (web, iOS, Android) into one hybrid codebase a smaller team can maintain. - **Weak security.** Common in vibe-coded apps. It looks perfect on the surface, but underneath no one can tell you how the data is actually protected. Open or wildcard permissions mean anyone who knows where to look can pull every user's personal data and wipe the whole database on the way out. That's not a bug you fix on Monday. That's the business, gone overnight. - **A limited talent pool.** The wrong core technology becomes a hiring handbrake. There are more than twenty times as many JavaScript developers as Elixir ones, so a niche stack chosen because a co-founder happened to know it can stall the whole org the moment you need to grow the team. Read these as measured, not felt. "The code feels bad" isn't a diagnosis. ## Start with discovery, then measure Before you pick a path, get honest about two things. First, the system: where it is today versus where it needs to get to. Second, the business: what you're trying to achieve, and what's actually in the way. There are always a hundred things wrong with a codebase. An honest discovery separates what blocks the business from what's just an engineer's nice-to-have. You fix what stands in the way of the goal, and the rest waits. Then put real numbers on it. Four things worth measuring: - **Technical health.** Response times, error rates, uptime. What your users actually feel. - **Complexity.** How many platforms, subsystems, screens, endpoints, and tables. The more moving parts, the more every change costs you. - **Quality.** Test coverage, security posture, how bad the debt is, whether the architecture can still hold up. - **Business value.** Time-to-market, cost to maintain, and the one everyone skips: which parts of the product people actually use. That last one tells you what's worth carrying forward and what's just dead weight you can drop. Now you're not guessing. You're choosing on evidence. ## The four paths Most people frame this as a choice between rewriting and leaving it alone. There are four options, and they line up like a ladder, from doing the least to doing the most: 1. **Continue as is.** Keep shipping on the system exactly as it stands. This fits only when the pain is low and nothing on the horizon needs more from it. The honest risk is that this is usually drift dressed up as a decision. And even if you stay put, don't deepen the debt. Hold every new feature to a higher standard, with its own tests, so the old debt stops growing and the parts you touch most often start getting better on their own. 2. **Stabilize.** Keep the system and clean it up gradually, to an acceptable shape rather than a perfect one. This is the right call when the bones are sound and the real problem is debt, coverage, and security drift, not the fundamental architecture. Pin the existing behavior down with tests first, then refactor safely behind them, starting with the lowest-hanging fruit. One healthcare platform we modernized this way went from effectively zero tests to over 95% coverage, moved off its end-of-life stack, and cleared an enterprise security review, all while it kept running. AI did much of the heavy lifting, which is what made catching up that fast affordable. 3. **Strangle incrementally.** Stand a new system up next to the old one and migrate feature by feature until the old one can be switched off. Martin Fowler named this pattern the Strangler Fig. It fits when you need a new shape but the business can't stop delivering while you get there. A decade-old media platform moving terabytes a month couldn't go offline, so we built the new platform alongside the original and rolled users over gradually, with no big-bang cutover. Downtime incidents dropped 90%, infrastructure cost dropped 80%, and the business ran the entire time. 4. **Big-bang.** Stop, rebuild, and cut over. It's the most debated path, and it fits only in a few narrow cases: when the foundation fundamentally can't carry the promise, when the system is small or early enough that you're not throwing much away, or when a hard external deadline leaves no room to be gradual. Before you pick it, answer Joel Spolsky's question honestly. What makes you think you'll do a better job the second time? One warning across all four. AI has made the incremental paths cheaper, and it's even made rewrites more reasonable than they used to be. But the safety doesn't come from the model. It comes from tests, guardrails, and rolling out in stages. When AI translates old code into new code, it makes mistakes, and it makes them systematically. Only a real test suite catches them. ## How to choose: three questions There's no scoring grid. Just three questions, asked in order: 1. **Does the foundation need to change at all?** If no, continue as is. 2. **Can the current foundation carry the promise?** If it just needs cleanup and the bones are good, stabilize. 3. **Can you get there while the system keeps running?** If yes, strangle. If you truly have to stop and switch, big-bang. The whole thing comes down to that third question: the path isn't set by how bad the code is. It's set by whether you can change it while it keeps running. ## The hard case: the staged conflict The hard projects have a staged sequence of goals that depend on each other, and not enough resources to modernize even once before the next milestone forces the issue. You win a deal on a patched-together, prototype-grade version. The client lands. Now there's finally money to do it right. But the bar keeps rising, from "make it demoable" to "make it actually work in production," while the budget only unlocks after the deal closes. So you owe a better system before you can pay for it. At bottom, this is a bet about timing. Do you invest ahead, spending to get the system ready before the contract is even signed, and speculate that the deal lands? Or do you win the deal first and take on the risk of catching up on the debt once the resources arrive? Neither is free. One bets your cash, the other bets your delivery. Just make that bet on purpose, not by accident. ## It's a phase, not a destiny Whichever path you pick, it isn't forever. The paths sequence. You continue as is until the debt starts to bite, then stabilize, then strangle when the promise needs a new shape, or jump straight to a rebuild when something forces it. Re-run the whole decision at every major milestone. The right answer this quarter isn't a promise you owe next year. ## Why now Every few years, software has to be redone. That's just the nature of it. Fifteen years ago everything moved from the desktop to the web. Ten years ago it moved to mobile. The shift now is toward systems that are agentic, ready for AI to operate them, and open enough to plug into everything around them. Each wave left a lot of perfectly good software a generation behind. What's different this time is the cost. AI has made modernizing and rebuilding genuinely affordable, because work that used to take years now takes months. So decisions that were quietly off the table, like the rewrite, the migration, or the clean architecture you could never justify, are back on it. But the honest first move is never to pick a path. It's to measure, so you choose on evidence instead of a gut feeling or a vendor's pitch. That's why we tell our clients to start with an audit before they commit to any path: an outside, senior look at what you actually own, what's breaking, and what to do first, with your own numbers, not ours. ### Why Software Keeps Breaking — and How It Should Be Built URL: https://brival.co/insights/2026/07/software-delivery-flow Category: Methodology · By Mat Kupczyk · 2026-07-05 Software rarely breaks because your developers are bad — it breaks because there’s no process around the code. This is the Software Delivery Flow: the seven stages, from strategy through monitoring, that turn “hand it to a developer and hope” into a product the business can actually see and steer. Most software doesn’t keep breaking because the developers are bad. It breaks because there’s no process around the code — and that process is the part most teams never run on purpose. The video above is the full walkthrough; this is the short written version, for skimming and reference. ## The problems we see over and over Writing code is the part everyone pictures, and it’s rarely where things go wrong. Shipping software is a much wider process, and it’s the part almost nobody runs deliberately. The pattern we see, especially at smaller companies, is always the same: the business hands an idea straight to a developer, and off they go. Then: - What comes back was **built to a guess** — the business finally sees it and says “that’s not what I meant,” and now you’re paying to redo it. - **Bugs reach customers instead of getting caught before release** — and they damage more than the code. - The **structure underneath gets you from A to B, but not to C** without tearing it up and starting over. - The whole process **lives in one person’s head** — so they become the bottleneck, and quality swings with whoever picked up the task. None of that is a coding problem. It’s the absence of a process. The odd part is that none of this is new. The way good software gets shipped hasn’t really changed in decades, and AI didn’t turn it upside down — if anything, it just made the discipline cheaper to reach. The shape has been well understood for years. Most teams still run it on instinct, and that’s exactly where it breaks. ## Our Software Delivery Flow We run every project through the same seven stages, and the last one loops back to the first. Here’s what each one makes sure of: 1. **Strategy** — everyone knows _why_ we’re building something, and who it’s for, before anyone starts. Put the goal and the audience in the ticket itself; if the why won’t fit in a sentence, it isn’t clear enough to build yet. 2. **Requirements** — that why becomes something the team can actually build, and the business can actually verify. Write down the expected behavior and how it changes what’s already there, and give one person the job of confirming that what shipped is what was asked for. 3. **Architecture** — the structure can grow without a rebuild, so one change doesn’t break five other things. Ask the hard questions about the future before you commit to a shape, let a senior make the call, and record the big decisions so the reasoning outlives whoever made it. 4. **Coding** — the work is written to one shared standard, so quality doesn’t swing with whoever picked up the task. Agree the conventions, where the code lives, and the test coverage each feature ships with — then add automatic checks so the standard doesn’t quietly erode under deadline pressure. 5. **QA** — problems get caught before your customers do. Bake testing into the process instead of leaving it to one person poking at the end, and every time you fix a bug, write the test that keeps it fixed. 6. **Release** — a change reaches production safely, predictably, and with everyone who needs to know in the loop. Keep releases small and frequent, automate the steps, and rehearse the rollback before you need it — an untested rollback is just a hope. 7. **Monitoring** — you know the system is up, the business is actually working, and whether the outcome you wanted moved — which feeds straight back into strategy. Watch two layers, technical health and business health, and block a recurring half-hour to actually read the numbers and act, rather than waiting for an incident. The point isn’t seven people, or a mountain of process. It’s that every one of these is _owned_ — even if, on a small team, one strong person owns several. What breaks a delivery is a stage left to no one. ## What actually makes it work Knowing the seven stages is the easy part. Whether they work for you comes down to a few things that have little to do with the diagram: - **It’s usually a clarity problem, not a talent problem.** Most of the messes we walk into aren’t bad work — they’re blurry or unspoken responsibilities. Name who owns each stage, and most of it resolves. - **Because it’s a process, you can measure it.** How many features get reworked, how many bugs reach customers, how many releases go messy. A handful of numbers, talked through with your team regularly, is most of the management — and it’s how a non-technical owner gets real oversight without reading code. - **Quality stopped being a budget question.** In the AI-assisted era, code is close to a commodity — so serious quality, full test coverage and an architecture you don’t have to apologize for, is a question of where you put your attention, not your budget. Owning the whole process end to end used to take a big team and a big budget; it doesn’t anymore. - **AI is a multiplier, and it works on whatever you feed it.** Give it vague requirements and no standard, and it produces low-value work five times faster. Give it real discipline — a clear why, a sound architecture, a standard it can read — and a single strong engineer can outperform a team of three or four, at a quality most companies couldn’t afford two years ago. The flow is what decides which way it multiplies. That’s what it looks like when software is built on purpose instead of on instinct: not handed to a developer and hoped over, but a process the business can actually see and steer. ### From Task Breakdown to Product Ownership: Project Management After AI URL: https://brival.co/insights/2026/07/project-management-after-ai Category: Lessons · By Mat Kupczyk · 2026-07-03 The AI conversation is stuck on speed. On our recent AI-first builds, the change that mattered more was to project management itself: the unit of work moved up a level, one developer now owns a whole product deliverable (data, API, UI, infra), and detailed planning moved into the middle of the build. Here’s what we’ve seen shift, and the one thing that didn’t. We’ve started a handful of greenfield projects over the last few months, all of them built with AI from the ground up. And the first thing that changed wasn’t the speed, it was the shape of the team. It’s smaller now, two people where we’d once have staffed five, and each person covers a lot more of the delivery than they used to. The same developer holds the architecture, builds it full-stack, and carries enough business understanding to know what the client actually needs. Fewer specialists handing work between them, more generalists who each own a deliverable outright, sitting close and talking often. The way we run those projects has drifted just as far, fewer calls and barely any ceremonies. But the more meaningful shift is structural. The unit of work a developer commits to moved up a level, and the detailed planning moved from _before the work_ into _the middle of it._ That quietly rearranges most of what our project management used to focus on. ## One developer now owns the whole screen A senior full-stack developer can now sit down in the morning and commit to having an entire screen finished by the end of the day: the data model behind it, the API endpoints, the UI, and the automated tests that keep it working. The commitment is to finish it, not just to start. A year ago that same screen was a whole feature story. We’d refine it together, split it along the seams of the system (someone takes the endpoint, someone takes the table, someone takes the frontend), and hand the pieces out. We broke it down because no single person could carry the whole thing inside a day, so we cut it into parts small enough to schedule and pass around.The ticket didn’t get faster, it got denser. A single developer-day now holds what we used to break into five subtasks across a few people.
AI absorbed the layer underneath. The SQL, the endpoint wiring, the boilerplate, the test scaffolding: all the stuff we used to decompose a story into is now the cheap part. It gets done in the same sitting as the thing it supports, so the sensible size of a ticket rose to match it. **What we track now looks a lot more like a product deliverable than an engineering task**, and one person owns it end to end instead of assembling it from parts owned by five. ## Detailed planning stopped happening before the work The old rhythm put planning first. You’d refine a story, argue out the approach, write the acceptance criteria, and only then would anyone open an editor. The plan was the contract, and the build was that plan carried out. That order has flipped, as now we talk about the product and the deliverable first: what this screen is for, what the customer will do with it. Then someone builds, and the plan firms up as the build shows us what the work really is. Planning didn’t disappear, the high-level product thinking still comes first. What moved is the detail: the cheapest place to work out the breakdown is now during the build, not before it. You prompt your way through an approach, and the model surfaces a decision you didn’t see coming (do this other thing first, then circle back).The real shape of the work only shows up in the doing. Trying to pin all of that down upfront in a perfect spec is slower, and less accurate, than just building a version and adjusting.
One of our developers put it plainly, from the management side. He wants one product story, defined well from the product side, and that’s the commitment he’s willing to make. He doesn’t want to spell out the ten subtasks underneath it in advance, and we agreed he shouldn’t. A breakdown written before the build is mostly guesswork, and the build outdates it within the day. The planning and the implementation have fused together, and prying them back apart just to produce a pre-agreed spec is busywork. ## Splitting the work costs more than it saves Here’s the part that’s easy to gloss over. Everything above makes a single capable person dramatically faster, but it doesn’t mean five of them ship five times as much. If anything, splitting one deliverable across two people now often costs more than just keeping it with one. To hand a slice of work to someone else, you need a clean seam between planning and implementation: a spec complete enough that another person can build against it without you. But the speed of the AI-first flow comes from fusing those two together. Pull them apart to parallelize, and you’re back to writing the exact upfront spec that AI let you skip. Then you hand it to someone who waits to be told what to do, instead of finding it in the build. You pay the old coordination tax right back, and get a more passive builder for it. It’s worst at the start of a project. When each person owns a whole screen, you’ve got two people making big changes at once. And there’s no solid core yet, no firm boundaries to keep them apart. Everything still touches everything, so the blast radius is high and their changes collide. **Once the architecture settles and the seams harden, parallel work gets safer.** Early on, a second person mostly just creates conflicts to untangle. Guardrails help, but only so far. We keep architectural oversight on the core decisions, and we lean on spec-driven development to hold the boundaries in check. Even then, the detail is where it bites: when each person’s PR solves its problem locally, inside its own module, the same logic ends up built in three places when it belonged in the core. Keeping that DRY across parallel work is troublesome. ## Smaller teams drift toward Kanban on their own So the teams got smaller - not in spite of the speed, but because of it. One or two people on a project is the comfortable default now, and the developers themselves push in that direction.“Fewer calls, drop the rituals, shorten the loop, let me just talk to the one other person on this build directly when something comes up.”
Below a certain size, all that coordination machinery is pure overhead. The whole point of process was to tame the entropy of a large group, and a group of two doesn’t have much of it to tame. So the pull is toward Kanban: short cycles, continuous flow, very little ceremony. The risk sits at the far end of that pull. Kanban with no committed horizon slides very easily into a black box, where work is happening and the developer is heads-down, but nobody outside can say where it stands or when it’ll land. That’s the exact thing we spent two decades building process to avoid, and moving fast isn’t a good reason to walk back into it. ## The layer that survived is the product layer If that’s the risk, what’s actually left to steer? A product backlog and an engineering breakdown were always two different things. The product layer (this screen, this flow, a checkout here) changes slowly and stays readable. The breakdown beneath it, the part that always bloated the board, is exactly what AI now handles inside the build. So what’s left to manage is the product layer, which was always the part that stayed manageable anyway. That’s why this doesn’t feel like project management getting harder. It gets narrower, and clearer: you steer what the client wants and whether it’s landing, not how the work breaks down under it. It also moves the scarce skill. The bottleneck isn’t the person who can slice a story into schedulable parts anymore. **It’s the person who can shape the deliverable from the product side, which is a lot closer to a product owner than a project manager.** The administrative half of the old role, tracking and distributing all those parts, is most of what AI took off our hands. ## Predictable delivery still needs an owner A couple of things didn’t change at all, and the speed only raises the stakes on them. The client still needs a commitment they can rely on. You can drop the word “sprint” and you can drop the ceremonies, but you can’t drop the predictable interval that someone owns: a window where the team commits to what’ll be done, ends it with something the client can actually see, and takes responsibility for the gap when it misses. The value of a sprint was never the ritual, it was the predictability and the reflection it forced. That survives AI completely intact. A client can’t hear “it’ll be done when it’s done” from us any more happily than we’d hear it from a contractor redoing our kitchen.You delegate responsibility for an outcome, not a decomposed list of tasks. The list is the cheap part now. Owning the result is the actual job.
And someone still has to hold the product clearly enough to define the deliverable and judge whether it landed. That was always the hard part. AI stripped away the busywork that used to disguise how hard it is. There’s less to manage now, but what’s left is the part that was never really administrative in the first place. One caveat before you run with any of this: we’ve only worked this way for a few months. Take it as a field report, not a finished method. A lot of teams are pushing the same pace right now, and I’d expect the playbook to look different a year from now. ### What AI Changes in a Software Estimate, and What It Leaves Alone URL: https://brival.co/insights/2026/06/what-ai-changes-in-a-software-estimate Category: Findings · By Mat Kupczyk · 2026-06-18 AI compresses how fast software gets built, but not the manual verification that confirms a high-stakes output is correct. The old “QA is a third of dev” rule would have badly under-budgeted it. Scope QA from the product, not from dev hours. AI compresses the work of writing code, and most of the work around it: documentation, test-case drafting, the operational setup. It does not compress the manual verification that confirms a high-stakes output is actually correct. So the estimate keeps its size in some lines and loses it in others, and the ratios we used to estimate by stop holding. On a project we just finished, development came in at 220 hours and quality assurance came in at around the same. A year ago the build alone would have run well over 400 hours by hand. And on the QA rule we used then, a third of the development we actually ran, we would have set aside _about 70 hours_ for it. This is not a one-off. We have shipped a handful of AI-built projects through 2026, and the same gap has shown up across them: _development keeps compressing, verification does not_. This is the one where we instrumented it closely enough to trust the numbers, and it changed how we scope the next one. ## The project This was a small project for a startup, promoting an early prototype to a proper-quality MVP. A greenfield build on Supabase with a React frontend, written AI-first from the start. It renders documents that have to match an exact format. When the format is wrong the document gets rejected downstream, and a rejected document is lost revenue, so “looks about right” was never going to be good enough. Speed did not come from cutting corners. Delivering it at roughly twice the pace of a hand-written build was not a matter of skipping the work that protects it: the build shipped with strong automated unit coverage on the backend and the infrastructure automation underneath it, more than a budget this size usually buys, because AI made both cheap to build. The point of saying so is that the QA load below is not a sloppy build catching up with us. It is the cost of verifying a correct build against a bar that does not move. ## The line that did not moveThe amount of product you have to check by hand is set by the product, not by how fast the product was built.
The number of formats to verify, the edge cases that change revenue, the real documents to run through the system: none of that got smaller because the code arrived sooner. It arrived sooner and then sat in front of the same verification it always needed. AI helped on either side of that verification, drafting and automating faster. What it did not do was make the judgment call at the critical path: does this output actually clear the real-world bar. **That call stayed manual, because the cost of being wrong is paid in revenue.** An automated check we have not yet confirmed against reality is not the thing standing between the client and a rejected document. ## The flip, in numbers For years we estimated QA at roughly a third of development time. It was a fair rule when development and verification scaled together. On this project the rule broke: - Old rule: QA ≈ 1/3 of dev. At 220 dev hours, that budgets about 70 hours of QA. - What happened: QA came in at 213 hours, almost exactly level with development. **The rule did not just become inaccurate. It failed in the dangerous direction.** It would have under-funded the one part of the work that got larger relative to everything else, on the project where being wrong was most expensive. If you anchor QA to dev hours and AI keeps halving dev hours, you will keep cutting the verification budget precisely as the amount to verify holds steady. ## QA signs off the behavior, then automation freezes it So why didn’t that cheap automation, the backend coverage and infrastructure work above, take the place of the manual QA? The answer is sequencing. An automated test freezes a behavior so the system cannot drift away from it later. **You can only freeze a behavior once someone has confirmed it is the right one.** During development that confirmation is a QA signing off on what correct looks like, against real documents, by hand. On a high-precision output, where the bar is an exact format a person downstream will accept or reject, that sign-off is the work itself, and the automation captures it rather than producing it. So the manual verification and the automation are two steps in order, not two ways of doing the same thing. The verification happens now, at full size, because it is the prerequisite. The automation pays off later, in maintenance: once the behavior is frozen, the tests hold it in place every time the system changes, and the infrastructure absorbs the change without hand-holding. On a greenfield build there is nothing to hold in place yet, which is why the payoff sits in the future, not in the first pass. They are two different lines in the estimate, and collapsing one into the other is how you under-fund the work. ## The constraint moved There is a second effect that does not show up in the totals. Development ran at AI pace, which meant **work arrived at QA faster than QA could clear it.** The bottleneck moved. For a decade the slow step was building the thing; here the slow step was confirming it, and the queue formed in front of QA instead of in front of the developers. ## How we scope it now Two assumptions in the baseline moved, and both are straightforward to price in. **We assume the build takes far fewer hours now, because AI compresses it.** And **we assume a higher quality standard inside the same budget**, because the automated coverage and infrastructure work that used to be the first line cut are now cheap enough to include by default. Less development, more quality built in, the same envelope. The change that is easy to miss is the QA line. **We scope verification from the product, not from the development estimate.** We list what has to be checked and how exactly, then size that directly.Development hours inform the schedule, but they no longer set the QA budget. That is the correction that keeps an estimate from quietly under-funding the work AI just made larger relative to the build.
Two disciplines did not change, the finding only raises the stakes on them. Automated coverage and manual verification were always separate lines, one a maintenance asset and one a present cost; now that AI lets us afford far more automation, keeping them apart matters more, not less. And precision was always priced by what a wrong output costs, not by how the code was built.