ios · tech — OCT 9, 2026
I Reverse Engineered a Popular iPad App With AI
I started with one IPA file and ended with a complete rebuild plan.
I wanted to understand how a popular iPad app worked.
Not just what buttons it had, but how its documents, core engine, file formats, tools and interface fitted together.
I didn't have the source code.
The obvious thing would have been to give the IPA to an AI tool and ask it to explain the app.
That wouldn't have worked.
An IPA contains a compiled program, resources and metadata. It doesn't contain the original Swift or Objective-C project.
So I had to turn it into useful evidence first.
Start with the IPA
An IPA is basically a ZIP file containing an iOS app bundle.
I passed the .ipa file to unzip, which produced the .app directory.
Inside that directory was the main executable, property lists, localised text, compiled asset catalogues, fonts, storyboards, NIBs and other resources.
I passed the main executable to file and otool. That established that it was an ARM64 Mach-O binary and showed its platform, SDK and encryption metadata.
The executable was already decrypted. I didn't bypass encryption or remove DRM.
I passed Info.plist and the extension property lists to plutil, and passed the app to codesign to inspect its entitlements.
This gave me the app's supported devices, operating-system requirements, permissions, file types and extensions before I looked at a single function.
Extract the easy information first
A lot of useful information in an iOS app isn't hidden inside machine code.
I passed every .strings, .stringsdict and .plist file to plutil and small Python scripts.
That turned the binary property lists into searchable JSON and text.
It exposed interface labels, error messages, settings, default values, supported file formats and feature names.
I also passed the executable to strings, nm and otool -ov.
Their output contained Objective-C selectors, class metadata, archive keys, file names and pieces of Swift runtime metadata.
This was enough to write a broad description of the app, but not enough to explain how it actually worked.
Inspect the app's own file formats
The app bundled examples of some of the formats it could read and write.
I passed its ZIP-based resource packages to unzip, then passed the binary property-list archives inside them to plutil and Python's plistlib.
That exposed the package layout, thumbnails, source images and named settings.
I did the same kind of inspection on bundled document resources, recording package members, archive classes, keys and image files.
This is much more reliable than guessing a format from file extensions or class names.
Recover the declarations
I then passed the main Mach-O executable, not the whole IPA, to a tool called ipsw.
Its class dumper produced Objective-C .h files containing classes, superclasses, protocols, properties, instance variables and method signatures.
Its Swift dumper produced a large text file containing Swift classes, structs, enums, protocols and surviving field names.
I fed those headers and Swift text files into a Python script that grouped them by subsystem.
At this point I knew the shape of the program, but method signatures still didn't explain what happened inside each method.
Decompile the executable
I imported the same ARM64 Mach-O executable into Ghidra.
Ghidra did not receive the IPA, storyboards, asset catalogues, fonts or resource files. Its input was the main executable.
I ran Ghidra's headless analyser, then used a custom Java post-processing script to export one .c file for every function.
These files were C-like pseudocode. They were not the original source code, and they couldn't be compiled back into the app.
Their value was that they exposed branches, constants, method calls, archive keys and the order that operations happened in.
Make the decompiled output readable
The first export still had tens of thousands of anonymous functions.
Objective-C message calls also appeared as generic function addresses, which made simple code look much harder than it was.
I wrote one Ghidra script that read the ARM64 Objective-C stubs and resolved the selector each one called.
I wrote another script that received a TSV file containing function addresses and Swift type names, then applied those names inside the Ghidra project.
Ghidra then re-exported the pseudocode with readable Objective-C calls and labelled Swift type accessors.
A Python script received the exported .c files, grouped anonymous functions near their likely owning class and produced a searchable TSV index containing each address, function name and folder.
It was still approximate, but it was now practical to investigate.
Don't ask one AI task to read everything
The decompiled output was millions of lines long.
Giving all of it to one AI task would have produced a shallow summary or exhausted its context before it got to the important parts.
I used a Python script to divide the class folders into feature-sized assignments.
Each assignment was a TSV file listing the folders it owned.
Each writing pass received:
its folder-assignment TSV, the relevant Ghidra .c files, Objective-C .h files, the Swift type dump, the extracted interface strings, the function index and the existing high-level overview.
It was told to describe behavior in plain English, record uncertainty and never paste the decompiled code into the specification.
Splitting the work this way produced detailed chapters for the core engine, documents, file formats, import and export, media handling, the editor, settings, data management and the app shell.
Recover the interface separately
Ghidra only analysed the executable.
It didn't explain the view hierarchy stored in compiled storyboards and NIB files.
I passed every .nib file and the NIB files inside .storyboardc directories to a custom Python NIBArchive parser.
That recovered archived classes, view hierarchies, frames, constraints, priorities, outlets, actions and prototype-cell geometry.
I passed each compiled Assets.car catalogue to Apple's assetutil.
That produced the logical asset names, device variants, scales and rendition metadata inside each catalogue.
The executable, the NIB archives and the asset catalogues all described different parts of the same interface. None of them was enough on its own.
Create the plain-English specification
By this point I had enough evidence to create the thing I wanted from the start: a plain-English specification of the app.
Each writing pass received the relevant headers, decompiled functions, interface strings, resource metadata and layout records. Its job was to explain the behavior without copying the pseudocode.
The result was a set of Markdown chapters covering each subsystem's behavior, data flow, file contracts, interface layout, defaults, edge cases and remaining unknowns, with pointers back to the evidence.
This wasn't a summary of the app. It was a reference detailed enough to guide a new implementation, while clearly marking anything the evidence couldn't prove.
Turn uncertainty into a list
Some things still couldn't be recovered confidently.
The wrong approach would have been to fill the gaps with plausible values and forget that they were guesses.
I searched every specification chapter for unresolved selectors, unknown file names, inferred defaults and behavior that required a real device.
Each finding went into a decision ledger with its evidence, implementation consequence and verification test.
Confirmed mistakes were marked for correction. Confirmed quirks were marked for replication. Missing evidence became a blocker instead of an invented answer.
Keep assets separate from behavior
Rebuilding how an interface behaves doesn't mean copying its icons, fonts, bundled content, example files or branding.
I enumerated the complete app bundle with rg --files -uu, because the extracted directory was ignored by normal project searches.
Direct files were classified by path, format, size and SHA-256 hash. ZIP-based resource packages were inspected for embedded images. Compiled asset catalogues were inventoried by their logical contents.
Those records became a separate asset specification describing what each asset was for and whether it should be replaced, recreated, generated, requested from the operating system or reimplemented.
The original assets can be useful in a private reference build, because they prove that the reconstructed loading and layout contracts work.
But the distributable build needs independently made or licensed replacements, plus an automated scan that fails if an original asset is accidentally included.
Turn the specification into small steps
A large specification is still not an implementation plan.
I gave the completed specification chapters, decision ledger, asset documentation, platform requirements and hardware requirements to one final planning pass.
It did not need the binary or the original assets.
Its job was to turn the evidence into small steps, where every step produced one visible or mechanically verifiable result.
Each step received an ID, exact specification references, dependencies, a minimal change, an automated or simulator check, a real-device check where necessary, completion criteria and a checkpoint.
Deterministic things like file codecs can be developed test-first. Things like hardware input, graphics output and performance need to be prototyped and measured first, then protected with regression tests after their behavior is understood.
The AI bill was about $400
Altogether, the analysis, specification writing, cross-checking and planning used about $400 worth of model tokens.
Now I can actually build it
So that was the point of all this.
I went from having an IPA file I couldn't really learn from, to having my own detailed, plain-English description of how the app works.
It doesn't give me the original source code, and I still have to build everything. But I now know what I need to build, how the parts fit together and which details still need checking.
Instead of guessing my way through it, I can work through the specification in small steps and build my own version.