Using Local AI to Speed Up Mobile Application Penetration Testing

Introduction
With the rapid rise of AI—and more recently, agentic AI—tasks that once required hours of manual effort can now be completed in a fraction of the time. Whether it's drafting emails, troubleshooting a washing machine, summarizing research papers, or generating code, AI has become an invaluable productivity tool that helps us work smarter rather than harder.
The same transformation is beginning to take place in cybersecurity, particularly in Mobile Application Penetration Testing (MAPT).
Why a Mobile App Pen Test is so Time Consuming

Figure 1. A mobile device is positioned beside a workstation running Burp Suite Professional during dynamic application testing. This setup allows the tester to route, capture, and inspect mobile network traffic while investigating authentication, authorization, and data-handling behavior.
A typical MAPT engagement involves much more than simply running automated scanners. Standard activities include:
Reverse engineering Android APKs or iOS applications
Performing Static and Dynamic Code Analysis
Reviewing source code or decompiled code
Analyzing manifests and configuration files
Tracing data flows
Identifying sensitive APIs
Bypassing SSL pinning and root/jailbreak detection
Intercepting network traffic
Test authentication and authorization
Manually verifying whether a suspected issue is actually exploitable.

Figure 2. Frida running the frida-multiple-unpinning script against an Android device to identify and override certificate-pinning implementations.
Data Overload
Many of these activities generate an overwhelming amount of information. A single application may contain tens of thousands of Java or Kotlin classes, hundreds of XML files, numerous third-party libraries, and thousands of lines of decompiled code. Sifting through this data manually is often one of the most time-consuming aspects of a mobile penetration test.

Figure 3. Decompiled code in MobSF, highlighting the volume of application data reviewed during static analysis.
Why Local AI Shines for Mobile App Pen Testing

Figure 4. A sanitized Burp Suite request and response pair showing how the mobile application exchanges authentication data with its backend API.
This is where local AI has significantly improved my workflow. The use of local AI allows for a more efficient analysis of the decompiled code all while maintaining the confidentiality of client data, as everything stays on our devices and never sent to companies like Anthropic or OpenAI. While the efficiency of using these local models allows for prioritizing areas of interest, summarizing large volumes of code, correlating findings from different tools and more, it's important that we not become fully reliant on AI and instead treat it as any other tool that needs professional analysis of the output.
In this blog series, we'll walk through how we built an AI-assisted MAPT workflow, the tools we use, where AI excels, where it still falls short, and come to a conclusion based off our findings.
Inside the Local AI Test Environment
The AI-assisted MAPT workflow runs on our dedicated AI supercomputer capable of running up to 200B models. We pair this hardware with opensource software that allows us to run these AI agents privately and securely on our device.
Specifications | ||
Model | ASUS GX10 | ![]() Figure 5. The ASUS GX10 local AI system used to process mobile application artifacts within the controlled test environment.![]() Figure 6. ASUS Ascent GX10 product identification label showing the model number, power specifications, wireless information, and regulatory certifications. |
Platform | NVIDIA GB10 Grace Blackwell Superchip | |
CPU | 20-core ARM | |
GPU | NVIDIA Blackwell | |
Unified Memory | 128 GB LPDDR5X | |
AI Performance | Up to 1 PFLOP FP4 | |
Model Capacity | Up to ~200B parameters, workload/model dependent | |
Networking | 10GbE + NVIDIA ConnectX-7 | |
Wireless | Wi-Fi 7 | |
OS | NVIDIA DGX OS | |
Power Rating | 48V DC, 5A (240 W max input) | |
Manufacturing | Made in Taiwan | |
Since one of the most time-consuming tasks of MAPT is code analysis, we decided to perform an experiment by comparing the results of what we normally use, MobSF + manual inspection, to our AI agent, and the results were impressive. One of the biggest advantages of using AI is its ability to reason with the information you give it, allowing for detailed generation of responses in a short timeframe. This reasoning is what allowed the model to produce not only similar results to what MobSF gave, but steps on how to exploit the vulnerabilities it found, recommend steps to patch the vulnerabilities, and show new findings that MobSF failed to see.
At first glance it seems AI is fully capable of doing what we do in a fraction of the time, but further examination of its output proves otherwise.
What AI Agents Excel at and Why Human Validation Still Matters
Upon further examination we noticed that although the agent was able to find vulnerabilities that MobSF didn't show and provide accurate findings, that doesn't necessarily mean that they were all true. When you take into account just how much information was given to the agent and what it gave as an output (description, evidence, exploitability, recommendation, etc.), there are a lot of factors that need to be checked and examined. This is where the drawback of AI comes into play, as although it gives all this information about the code, if we were to blindly put our trust into its analysis or allow it to act upon its recommendations we would quickly see where AI starts to reach its limits. This is why it is crucial for processionals in the field to examine what these models are outputting in order to filter out what is true and what isn't.
Verdict: AI vs. Human
As this experiment shows, local AI has real potential to reshape how mobile application penetration testing is conducted, cutting down the time spent sifting through decompiled code, XML files, and third-party libraries while still surfacing findings that traditional tools like MobSF can miss. But speed and volume of output are not the same as accuracy, and this workflow only works because a skilled analyst is still validating every finding, checking exploitability, and discarding false positives before anything reaches a client's report. AI should be viewed as a force multiplier that handles the tedious groundwork of correlation and summarization, not as a replacement for the judgment, context, and hands-on verification that experienced testers bring to an engagement. Running these models locally also ensures that this productivity gain doesn't come at the cost of client confidentiality, since sensitive code and findings never leave our own infrastructure.

Polito Inc. offers a wide range of security consulting services including penetration testing, vulnerability assessments, red team assessments, incident response, digital forensics, threat hunting, and more. If your business or your clients have any cybersecurity needs, contact our experts and experience what Masterful Cyber Security is all about.
Phone: 571-969-7039
E-mail: info@politoinc.com
Website: politoinc.com






Comments