Case study · Qualgent.ai

YC X25 · SaaS & Technology · Dedicated Teams

AI agents that test mobile apps like human QA

Qualgent runs AI agents on real iOS and Android devices. Tests are written in plain English; self-learning agents adapt to UI changes and run in parallel across device farms.

PythonReactNode.jsComputer VisionLLMsWebSockets
  • 60×

    Faster regression testing

  • 90%

    Less test maintenance

  • #3

    App Store rank reached by a customer

  • 5

    Core systems shipped

The client

AI agents for mobile QA

Qualgent.ai uses AI agents to mimic real human testers on actual devices. No scripts, no selectors, no maintenance overhead — and parallel execution across device farms turns a three-day regression into minutes.

The problem

QA became the bottleneck of AI-accelerated development

As AI makes coding faster, teams ship code faster than they can test it. Manual mobile QA does not scale.

Script-based tools like Appium are flaky and break on every UI change, creating a maintenance loop that defeats the purpose of automation.

  • Manual QA cannot keep pace with AI-accelerated development
  • Appium / XCTest scripts break on every UI change
  • Maintenance cost exceeded the value of automation
  • No parallel execution across real devices
  • 3-day regression cycles blocking weekly releases
What we built

5 systems, one platform

Violet marks the parts where the AI does the work. A human approves every consequential change.

  1. Module 01

    Cloud device management

    Provisioning, health monitoring, session management and auto-recovery for real iOS and Android devices in the cloud.

    • Real device provisioning
    • Parallel execution across farms
    • Health monitoring & recovery
    • Session & resource allocation
  2. Module 02

    Web dashboard & test management

    Author tests in plain English, organise suites, schedule runs and review results with video and step logs.

    • Plain-English authoring
    • Suites & scheduling
    • Video playback per run
    • Step-by-step logs
  3. Module 03

    Integration layer

    GitHub, Slack and Linear integrations plus developer APIs so tests run on every build and post results to PRs.

    • GitHub, Slack, Linear
    • CI/CD developer APIs
    • Webhooks
    • Trigger on push
  4. Module 04 AI does this

    AI agent orchestration

    Computer-use LLM agents and vision models tap, swipe, type and speak like a human tester — and self-heal when the UI changes.

    • Vision-driven device interaction
    • Self-healing tests
    • Natural-language interpretation
    • Multi-step context
  5. Module 05

    Reporting pipeline

    Video, step logs and pass/fail evidence for every run, with failure categorisation and trends.

    • Full video per run
    • Pass/fail with evidence
    • Failure categorisation
    • Exportable reports
Stack
PythonReactNode.jsComputer VisionLLMsWebSocketsDockerCloud infrastructure
Outcome

60× faster regression testing.

More proof

Other systems we've shipped

Case study
YC S24AI Solutions

The compliance platform for safety-critical hardware

Saphira.ai replaces IBM DOORS-era tooling for robotics, automotive and aerospace teams. We built the core platform: AI document parsing, automated risk assessment and full requirement-to-test traceability.

10×

Faster compliance cycles

Read the case
All case studies

Have a similar problem?

Tell us about it. A senior engineer replies within 24 hours with scoping questions and a real timeline.

  • 50+

    Production systems shipped

  • 97%

    Client retention

  • 4–8 wk

    Typical time to production

  • 24h

    Reply from a senior engineer