DexShield · Benchmarks

DexShield benchmarks

What DexShield actually costs, and what it actually buys you. Every number below is measured, not estimated — and every one is reproducible with the scripts in the repository. Created by Ivan Garibay.

1. Size & speed (protection pipeline)

Corpus: 60 synthetic classes with sensitive constants (API keys, salts, license and device-binding strings), compiled with javac + d8 and run through the full protect-apk transform chain (class / member / virtual renaming, AES-256 DEX string encryption, strip-debug).

+29.9%
DEX size overhead
67.6 KB → 87.8 KB
1.29 s
Total protect time
≈ 21 ms / class
295 → 0
Sensitive strings
left in cleartext
584
Identifiers renamed
(mapping.txt lines)

Detailed measurements

MetricPlainProtectedΔ
classes.dex size67,584 B87,764 B+29.9%
Packed APK size11,480 B20,307 B+76.9%
Sensitive strings in cleartext2950−100%
Protection wall-clock time—1,290 ms—

The packed-APK delta is larger than the DEX delta because AES-256 ciphertext is Base64 and high-entropy, so it compresses far worse than the plaintext constants it replaces. In a real multi-MB APK — where DEX is a fraction of the whole and most strings are not encrypted — the relative overhead is much smaller; this corpus is deliberately string-dense to stress the string-encryption path. Every constant that carried meaning to an attacker is gone from the static image, at a cost of ~21 ms per class.

Visual

DEX plain67.6 KB
DEX protected87.8 KB
Cleartext secrets — plain295
Cleartext secrets — protected0

2. Reverse-engineering resistance (LLM-scored)

A different question: not how big, but how hard to understand. We give a state-of-the-art LLM exactly what a human reverse engineer gets — the disassembly, no source, no mapping — and score how much it can recover from the plain build versus the DexShield-protected build. Recovery is scored 0–1 across three tasks, over a corpus of 8 samples across 8 domains (licensing, crypto, networking, JWT auth, password hashing, feature gating, key-wrapping, root detection).

Objective, rater-free signals (8 samples)

Signal (per sample, plain → protected)PlainProtected
Secret visible in disassembly (baksmali/dexlib2)8/80/8
Intent-carrying labels surviving (class/method/field)32/320/32
Secret recoverable in raw string-pool (strings classes.dex)8/80/8

The bigger corpus caught a real bug. The one-sample pilot claimed the secret was absent from the protected DEX — but that was measured against the disassembly. Across 8 samples the raw string-pool check showed the plaintext still recoverable with strings in 7/8 cases: D8 keeps the default value of a static final String in the DEX static_values array, which the encryptor was not rewriting. That was fixed — those initializers are now stripped and re-assigned encrypted in <clinit> — so the plaintext is now gone from the raw image in 8/8. This is exactly the kind of weakness a single sample hides and a corpus exposes.

Recovery scoring (post-fix)

TaskPlainProtected
T1 — Semantics (what does it do?)1.00.6
T2 — Identifiers (recover real names)1.00.3
T3 — Secret (extract the salt)1.00.1
Average recovery1.000.33
0.67
Resistance = 1 − (0.33 / 1.00)
100% → 33%
LLM recovery, plain → protected
absent
Plaintext secret in DEX
(grep = 0)

The low-level algorithm shape (a ×31 rolling hash / XOR + a hidden constant + hex compare) survives — static protection does not hide control-flow arithmetic. What is destroyed is meaning: intent labels like LicenseValidator, isPremiumUnlocked, SECRET_SALT are unrecoverable (0/32), and the secret itself is no longer anywhere in the file. Combined with per-build diversification (the mapping differs on every build), an LLM-derived deobfuscation of one build does not transfer to the next.

Recovery scoring (T1/T2) is single-rater; the secret/label columns above are fully reproducible and rater-free. See RESULTS-CORPUS.md for the 8-sample run and RESULTS.md for the original pilot.

Diversification — one build's deobfuscation doesn't transfer

The same corpus, protected twice with different seeds, then measured for reuse ([bench/diversification]):

What an attacker reuses from build 1 against build 2TransferIdeal
Obfuscated names (same symbol → same name)1.2%0%
String ciphertexts (same key)0.0%0%
Decrypt runtime name/locationdiffersdiffers

Fully deobfuscating one build yields almost nothing against the next: 98.8% of names don't line up, not a single ciphertext matches (the AES string key is derived per build), and the decrypt runtime has moved — there is no fixed dexrt.S->d anchor to hook across apps. Measured, not asserted.

3. DexShield vs DexGuard

Feature-for-feature, DexShield covers the core of what the commercial reference tool does — including R8-compatible method virtualization (int/long, arrays, control flow, calls, constructors and fields, verified booting on a real Verifone T650p). DexGuard's virtualization is still more mature and broad. Full side-by-side table:

→ DexShield vs DexGuard: the full comparison

Reproduce it yourself

Nothing here is a marketing figure. Both benchmarks are scripts in the repo:

# Size & speed
JDK17="/path/to/jdk-17" bash bench/perf/run.sh
# → bench/perf/out/perf.json

# LLM reverse-engineering resistance
bash bench/llm-resistance/run.sh
# → feed bench/llm-resistance/out/protected.smali.txt to an LLM
#   with the tasks in bench/llm-resistance/README.md
⭐ DexShield on GitHub