Article

Base44 vs Emergent vs Lovable vs Replit: A Dated Test

A January 2026 hands-on test of Base44, Emergent, Lovable, Replit, and Firebase Studio using the same 122-page insurance app brief.

Aug 4, 20268 min readBy Dalton Anderson
In this article

Base44, Emergent, Lovable, and Replit on the Same App Brief

In a hands-on test recorded January 5, 2026, Base44 produced the strongest initial insurance-app demo of the tools Dalton Anderson tested. Lovable was the most pleasant surprise. Replit eventually produced a polished, interactive result but created more startup friction. Emergent consumed the available free credits while returning an experience that behaved mostly like a landing page. Firebase Studio generated interfaces, but it did not satisfy the intended one-shot test and should not have been called Genkit.

That is a dated lab observation, not a current ranking. The builders, plans, models, and interfaces have changed. More importantly, the test preserved the transcript and screen recordings but not the original 122-page requirements file, exact account tiers, model selections, receipts, or complete run logs. The result can show what happened. It cannot show how often the same result would happen again.

The input was one instruction, not one sentence

Dalton used a long insurance product document that described the intended product, JSON schemas, APIs, sprints, and implementation ideas. In the episode, he says it ran to roughly 122 pages and had taken a long time to assemble.

The test called each initial build "one shot" because Dalton gave the builder its instructions and then judged the first substantial output without coaching every intermediate step. Some systems asked clarifying questions. That distinction matters. One-shot describes the number of builder interventions before the first result, not the amount of human work behind the request.

The task was also unusually broad. It asked the tools to approximate a property-insurance workflow with submissions, property data, risk information, forms, carrier options, and dashboards. A builder that returned less could have failed, exhausted its budget, or interpreted the requested scope differently. The missing original brief prevents a line-by-line requirements audit.

What the evidence contains

The preserved package includes the raw transcript and two screen recordings captured on January 5, 2026. The recordings total about 62 minutes. They show live builder interfaces and generated previews for Base44, Emergent, Lovable, Replit, and Firebase Studio.

The method did not preserve a standardized timer, plan manifest, model manifest, spend ledger, clean-account state, retry policy, or multiple runs. Vendor comparisons should be read with that limitation beside every result.

Platform as testedVisible result in the recordingImportant limitation
Base44Multi-screen dashboard, submissions, properties, claims, analytics, and new-submission flowGenerated workflow and data were not verified as correct or production ready
EmergentPolished product page with a limited preview and generated performance claimsAvailable credits were exhausted and the result did not cover the requested breadth
LovableInteractive quote and bind-style demo with a distinct visual treatmentCarrier data and transactions were simulated
ReplitWorking submission and classification screens after substantial frictionTranscript records authentication, 404, and broken analytics behavior
Firebase StudioSeveral generated dashboards and app screensRequired more setup, showed instability, and the tested interface is now historical

Base44 made the strongest first impression

The Base44 recording shows more than a screenshot. The preview contains a dashboard, recent submissions, portfolio health, property records, claims, analytics, and a new-submission path. Dalton can move among screens and interact with the generated workflow.

Base44 risk dashboard captured during the January 5, 2026 E098 test

The image supports a narrow claim: Base44 generated a coherent, navigable risk-management demo in this run. It does not establish that the submission flow follows a carrier's process, that its insurance fields are accurate, or that authorization and persistence work.

Dalton also noted that the builder asked questions before building. That helped it narrow a very large request. The transcript and editor show at least one later code correction involving component import paths, so "one shot" should not be interpreted as error-free generation.

Base44's current documentation has moved well beyond the exact surface in the recording. Its app-building guide now describes managed data, authentication, permissions, testing, hosting, and publishing. Its GitHub documentation describes two-way synchronization on eligible plans, but also documents permanent connection and history constraints. Those current capabilities do not retroactively strengthen the January result.

Emergent looked polished before the application did

Dalton liked Emergent's builder interface and the thoughtfulness of its questions. The disappointment came from the ratio of time and credits to usable output. He recorded a negative free-credit balance and described the preview as a landing page rather than a working application.

The recovered frame shows why that distinction matters. The page presents claims such as 95 percent time saved, 100 percent data accuracy, and enrichment in under three seconds.

Emergent landing page with generated performance claims captured during the January 5, 2026 E098 test

Those figures are generated demo copy. They are not measurements. A credible review cannot repeat them as product performance.

The outcome may also reflect the test boundary. The request was large, the account was on a free tier, and the transcript suggests the system may have stopped when credits ran out. That makes the result useful as a buyer-experience observation but weak as a claim about what Emergent usually produces with sufficient budget.

Emergent's current product documentation describes full-stack generation, preview, testing, deployment, and integrations. Its GitHub guide documents push and pull workflows. Current plan and credit details remain too changeable to borrow for a permanent comparison.

Lovable was the strongest surprise

Lovable began building with less questioning and produced an interactive demo that covered more of the visible product story. The recording shows a quote sequence with carrier options, prices, match percentages, and a bind action.

Lovable interactive quote flow captured during the January 5, 2026 E098 test

The interface was not proof of real quoting or binding. It was evidence that the builder translated the brief into a broader clickable narrative. Dalton also preferred its styling and the way the editor let him observe the build.

Lovable's current getting-started documentation describes Agent and Plan workflows, managed backends, GitHub, and publishing. Its deployment and ownership guidance distinguishes managed, hybrid, and self-managed routes. That lifecycle detail is more useful for a present-day buying decision than a remembered January ranking.

Replit eventually worked, but the route was rougher

The transcript repeatedly says "Riplet" or "Ripplet." The visible product and context identify Replit. Public references should correct the name without preserving the transcription error as if it were a separate vendor.

Dalton spent about 30 minutes dealing with authentication and a 404 before reaching a useful application. Once it loaded, the preview looked professional and included a submission flow. The recording also shows generated risk-classification results.

Replit risk-classification flow captured during the January 5, 2026 E098 test

The app later broke when Dalton opened analytics. That is exactly the kind of evidence a first-screen comparison misses. Replit's output could be visually strong and still fail a navigation test.

Replit's current Agent guidance does not promise that the first output should be accepted. It tells users to be specific, plan, add context, review and test, and use checkpoints. Its current publishing documentation also distinguishes deployment types and warns against treating the published filesystem as durable storage.

The Google product was Firebase Studio, not Genkit

Dalton called the Google experience Genkit. The recording shows Firebase Studio's App Prototyping agent. Google describes Genkit as an open-source framework for building AI-powered and agentic applications. It is a developer framework, not the hosted prompt-to-app product in this comparison.

The recording shows Firebase Studio generating a dashboard with address ingestion, classification, document uploads, carrier matching, and portfolio views. It also shows a Gemini API-key prompt and a security-check warning.

Firebase Studio generated dashboard captured during the January 5, 2026 E098 test

Dalton found the result unstable and difficult to navigate. He also objected to the system making architecture choices without a richer requirements conversation.

The product context has since changed decisively. Google's Firebase Studio documentation says creation of new App Prototyping agent workspaces was disabled on June 22, 2026 and recommends Google AI Studio. It also warns that generated output can be inaccurate and that untested generated code should not be used in production.

Base44 won this run, not the category forever

Base44 was Dalton's clear favorite on the recording date because it combined speed, navigability, visual coherence, and more of the requested workflow. Lovable delivered the most notable improvement relative to his prior expectations. Replit produced a credible demo but lost ground through friction and broken behavior. Emergent's result was not worth the time and credit spend in this account. Firebase Studio did not fit the one-shot comparison cleanly.

None of those observations answers which builder is best now. A present-day decision needs the same task, documented plans and models, equivalent budgets, clean accounts, fixed stop rules, functional tests, domain review, security checks, portability tests, and multiple runs.

The more durable result from E098 is not the winner. It is the difference between a convincing first screen and evidence that a product works. The next useful step is [[How to Benchmark Vibe Coding Tools]], followed by [[From AI Prototype to Production]] when a generated demo is worth keeping.

AI assisted with research organization, structure, drafting, and validation. Dalton Anderson remains the attributed author and final editorial authority. The transcript and linked public sources control factual claims. Publication remains unauthorized.

Sources

Follow the evidence.

  1. docs.base44.com: githubdocs.base44.com
  2. web.dev: vitalsweb.dev
  3. docs.replit.com: replit appsdocs.replit.com
  4. docs.replit.com: build with agentdocs.replit.com
  5. csrc.nist.gov: finalcsrc.nist.gov
  6. help.emergent.sh: 272715 features and toolshelp.emergent.sh
  7. firebase.google.com: migrating projectfirebase.google.com
  8. owasp.org: www project application security verification standardowasp.org
  9. docs.base44.com: Quick start guidedocs.base44.com
  10. w3.org: WCAG22w3.org
  11. firebase.google.com: get started aifirebase.google.com
  12. help.emergent.sh: plans and creditshelp.emergent.sh
  13. docs.lovable.dev: githubdocs.lovable.dev
  14. docs.lovable.dev: getting starteddocs.lovable.dev
  15. firebase.google.com: overviewfirebase.google.com

From this episode

Two useful next steps.

Evergreen · 1 min

What One-Shot App Generation Actually Proves

A one-shot AI app build can prove initial instruction-following and visible interaction. It cannot prove security, correctness, scale, or demand.

Research Note · 1 min

Vibe Coding Benchmark Method Research Note

A useful AI app-builder benchmark must answer a decision rather than manufacture a universal leaderboard. The decision might be which tool best supports a team's internal

Return to the episode