Conclusions and decision conditions
- First, inventory models, inference runtimes, pre/post-processing logic, prompt templates, API credentials, caches, logs, and server-side resources separately.
- File encryption protects the static storage phase; however, the runtime state after loading, API invocations, and outputs require separate controls.
- Long-lived high-privilege API Keys should not reside as client secrets. Prioritize server-side proxies, short-lived tokens, least privilege, and revocable strategies.
- Play Integrity and App Attest provide evidence of application instances or environments, but final authorization and risk remediation remain on the server.
Deconstruct Mobile AI Applications into Seven Asset Classes
Mobile AI security is often oversimplified to whether the model is encrypted. In reality, the attack surface includes at least model files, inference runtimes, input preprocessing, output post-processing, prompts or business rules, remote API credentials, user data, and logs. Each asset class carries different consequences upon leakage, distinct update mechanisms, and separate ownership responsibilities.
Copying model weights may lead to intellectual property theft and loss of business capability; modifying prompt templates or post-processing logic can alter business behavior; leaking long-lived API Keys can directly result in financial loss, unauthorized data access, or resource abuse; input/output logs may contain user privacy data. Protecting only the model file fails to cover these remaining risks.
An asset register must document storage location, generation source, transmission path, runtime usage patterns, update and rollback mechanisms, least privilege principles, log scope, and retention periods. Assets without a defined lifecycle should not be deemed protected solely by enabling an encryption switch.
| Asset | Primary Risk | Priority Control | Cannot Be Replaced By |
|---|---|---|---|
| Model Files | Copying, static analysis, version substitution | Controlled delivery, file protection, integrity checks, and version pairing | API Authorization |
| Inference Runtime | Injection, debugging, memory inspection, dependency risks | Application hardening, dependency governance, anomaly and compatibility validation | Encrypting model files only |
| Pre/Post-Processing Logic | Rule reversal or modification | Critical path protection, server-side verification, and regression testing | Model weight protection |
| API Credentials | Abuse, cost overruns, and data privilege escalation | Server-side proxy, short-lived tokens, privilege restriction, and rotation | Code obfuscation or model encryption |
| User Input/Output | Privacy leakage, injection attacks, sensitive data echo | Minimal collection, filtering, masking, and access control | Generic promises of device encryption |
| Logs and Caches | Long-term retention of sensitive material | Classification, masking, expiration, and controlled export | Disabling a single debug switch |
| Server-Side Resources | Unauthorized invocations and automated abuse | Account policies, quotas, behavioral analysis, versioning, and integrity strategies | Client-side self-reported trust |
Model Files Must Become Usable Material During Runtime
Whether using Core ML, LiteRT, or other on-device runtimes, models must be loaded, parsed, and engaged in computation. File-layer encryption reduces the convenience of direct copying from installation packages or app directories, but it cannot prevent the running application from accessing model materials. If an attacker controls the process or observes the runtime, the risk model shifts from static files to loading, memory, invocations, and outputs.
This does not imply model encryption lacks value. It raises the cost of low-effort copying, prevents direct substitution, and forms part of a defense-in-depth strategy alongside package protection, integrity checks, runtime environment detection, and version pairing. The key is to frame the benefit as increasing attacker cost and reducing the exposure surface, rather than claiming extraction is impossible.
Model updates also introduce compatibility challenges. Model versions must be paired with preprocessing logic, feature shapes, runtime versions, hardware capabilities, and post-processing rules. Update failures require safe rollbacks to prevent random combinations of old models and new logic.
| Phase | Exposure Surface | Control Focus | Validation Questions |
|---|---|---|---|
| Build and Packaging | Repositories, CI pipelines, installation packages, and resource directories | Access control, key isolation, file protection, and artifact identity | Which models and configurations appear in the final package? |
| Download and Update | Network, cache, and version switching | Transmission protection, signing or integrity checks, atomic replacement, and rollback | Will anomalous updates load mismatched models? |
| Loading and Inference | Process memory, runtime interfaces, and hardware backends | Runtime protection, minimal residency, exception handling, and compatibility | When does the model become available, and how does failure halt execution? |
| Output and Logging | Results, confidence scores, debug information, and user data | Minimal output, masking, access control, and expiration | Do logs leak model details or sensitive user information? |
Long-Lived High-Privilege API Keys Should Not Be Client Secrets
Android security documentation explicitly states that API Keys embedded in source code may be discovered via decompilation after the app is compiled. Obfuscation and application hardening can increase the cost of locating them, but they cannot change the fact that the client must hold and use these credentials. As long as a key remains long-lived and highly privileged within a generic client, the impact of leakage is difficult to constrain.
Android Keystore can keep certain key materials non-exportable, but official documentation clarifies that if the application process is compromised, attackers may still be able to use the app's keys to perform operations. It is suitable for protecting device-bound private keys and local encryption, but should not be misinterpreted as a secure vault for arbitrary long-lived shared secrets for remote services.
A more robust architecture requires the client to request short-lived, restricted, and revocable tokens from its own service based on user identity and device context, or have the server-side proxy handle high-risk model invocations. The server must enforce account authorization, quotas, rate limiting, model scope, data scope, and anomalous behavior monitoring.
- Isolate credentials by development, testing, and production environments
- Ensure tokens possess minimal model and data permissions
- Enable server-side revocation and enforce quota and rate limits
- Prevent clients from storing long-lived high-privilege shared keys
request = verify_user_session(input.session)
app = verify_app_evidence(input.attestation)
policy = load_policy(user=request.user, app=app.identity)
if policy.version_state == UNKNOWN:
return CHALLENGE_OR_LIMIT
if policy.account_scope.allows(input.model_scope) == false:
return DENY
if policy.risk_score >= HIGH:
return STEP_UP_VERIFICATION
return issue_short_lived_token(
scope=input.model_scope,
quota=policy.quota,
expires_in=policy.short_window
)Integrity Signals Should Only Inform Server-Side Decisions
Play Integrity provides Android backends with signals regarding app identification, device integrity, account licensing, and partial environmental risks. Apple App Attest helps the server determine if a request originates from a valid app instance by generating device keys, issuing one-time challenges, verifying attestation on the server, and processing subsequent assertions. The commonality is that evidence is ultimately validated by the server.
These mechanisms have clear boundaries. Play-related signals are influenced by distribution sources, device status, and service conditions; Apple documentation notes that App Attest is not supported on all device types and that no single policy eliminates all fraud. The server must distinguish between pass, fail, unavailable, transient errors, and unconfigured states. It should not equate 'unavailable' directly with an attack, nor should it perform final judgments locally on the client.
Integrity signals and model encryption address different problems. The former helps verify application instances and runtime environments, while the latter reduces static model exposure. True authorization still requires account context, business resource checks, version validation, request content analysis, quota enforcement, and behavioral context.
| Layer | What It Provides | Who Validates | Common Misuse |
|---|---|---|---|
| Model File Protection | Resistance to static storage access and substitution | Client and delivery chain | Treating file encryption as API authorization |
| Play Integrity | Signals related to Android apps, devices, accounts, and environment | Server-side | Client returning a trusted boolean value itself |
| App Attest | Attestation and assertion for Apple app instance keys | Server-side | Ignoring challenges, counters, or unsupported devices |
| Account and Business Policies | Whether a user can access specific models, data, and quotas | Server-side | Relying solely on device signals while ignoring user permissions |
| Behavioral Risk Control | Rate limiting, replay detection, bulk abuse prevention, and anomalous context | Server-side | Granting permanent trust after a single pass |
Validating Mobile AI Applications Requires Static, Runtime, and Server-Side Views
Static checks answer what models, configurations, strings, credentials, and debug resources exist in the installation package; runtime checks answer how models are loaded, how failures are handled, whether logs leak data, and if updates can roll back; server-side checks answer whether account, version, integrity, quota, and data permissions are truly enforced. Missing any of these three evidence types biases conclusions toward partial views.
Testing must use unique release candidates and explicit model versions, documenting target systems, device capabilities, runtime backends, model sources, and update states. On-device inference performance, memory usage, and compatibility depend on model structure, quantization, hardware, and runtime; figures from other models or devices cannot be cited.
Exception paths are critical: missing or corrupted model files, interrupted updates, unsupported runtimes, expired server tokens, unavailable integrity signals, failed log uploads, and revoked user permissions. The system must degrade safely, providing recoverable states for both users and operations teams.
| Evidence Surface | Check Content | Pass Conditions | Conclusion Boundaries |
|---|---|---|---|
| Installation Package (Static) | Models, keys, configurations, log markers, and debug resources | No long-lived high-privilege secrets; model delivery aligns with design | Cannot prove runtime is unobservable |
| Model Runtime | Loading, memory lifecycle, errors, performance, and rollback | Stable on target devices with safe exception handling | Cannot prove server-side authorization is correct |
| Network and Credentials | Token validity, permissions, rotation, replay protection, and certificate policies | Least privilege and revocability | Cannot prove model files are protected |
| Server-Side Policies | Accounts, versions, integrity, quotas, and behavior | High-risk resources determined finally by the server | Cannot treat a single platform signal as absolutely trusted |
| Privacy and Logs | Input/output, caches, diagnostics, and exports | Minimal collection, masking, and expirability | Must combine with business data classification requirements |
Design Conclusions and Scope Limitations
Model encryption is worthwhile, but it represents only one control point within the model lifecycle. Commercial AI apps also need to protect pre/post-processing logic, prevent long-lived high-privilege keys from entering the client, ensure the server performs final authorization, implement fault tolerance for platform integrity signals, and govern user data and logs.
Without actual release candidates, model versions, target devices, and server-side interfaces, we can only review architecture and pending validation items. We cannot claim models are extraction-proof, interfaces are abuse-proof, runtime injection is blocked, or performance targets are met. Any such conclusions require current evidence packages.
The action entry point for Yudun on this page is for applying for application hardening and compatibility assessment; it does not represent validation of any specific model framework or AI-exclusive capability. Project scope must be confirmed separately after submitting the tech stack, model delivery method, and critical business paths.
- Maintain separate registers for models, runtimes, credentials, and data
- Ensure high-risk invocations are finally authorized by the server
- Provide unavailable and degradation paths for platform signals
- Enable model updates to support verification, atomic switching, and rollback
- Bind all conclusions to specific release candidates and model versions
Evidence and applicability boundaries
This section separates documented platform facts, engineering judgment, and limits that cannot be generalized into unverified product claims.
| Article judgment | Fact or engineering basis | Applicability limit |
|---|---|---|
| API Keys compiled into the client may be discovered via decompilation. | The official Android Security Checklist explicitly states that when source code contains API Keys, attackers may decompile the app and locate these resources. | Some platform-restricted low-privilege Keys may reside on the client per vendor rules, but scope restriction and monitoring remain necessary. |
| Android Keystore cannot prevent a compromised process from using keys. | Official Android documentation states that while key material can remain non-exportable, attackers may still use the app's keys if the application process is compromised. | Specific protection capabilities depend on key usage, hardware support, authentication constraints, and implementation details. |
| App Attest attestations and assertions must be validated on the server. | The official Apple workflow uses server-side challenges, attestation verification, public key storage, and subsequent assertion counters. | Not all device types are supported, and Apple explicitly states that no single policy eliminates all fraud. |
| Model file protection cannot replace server-side resource authorization. | The two protect different objects: one targets client-side files and runtime materials, the other targets account, model, data, and quota permissions. | This is an architectural responsibility judgment and does not imply any specific model encryption implementation has passed validation. |
| Conclusions on on-device model performance and compatibility must be bound to specific models and devices. | Model structure, quantization method, runtime, hardware backend, system version, and input scale jointly influence results. | This article provides or implies no performance figures for Yudun models. |
Engineering questions
The model is already encrypted; why can't I still put the API Key in the app?
Model encryption protects model files, whereas API Keys are credentials for accessing remote resources. When the client must use the Key, attackers may still obtain or abuse it through static or runtime observation.
Is storing the API Key in Android Keystore absolutely secure?
Keystore can reduce the risk of key material being exported, but a compromised application process may still invoke the key. It is better suited for device-bound operations and does not replace server-side least privilege and short-lived tokens.
Can devices be permanently trusted after passing Play Integrity or App Attest?
No. They provide evidence for a specific time and context. The server must still validate accounts, requests, versions, quotas, and behavior, while handling signal unavailability and state changes.
Must on-device models be migrated to the server?
Not necessarily. Offline, privacy-sensitive, and low-latency scenarios may require on-device inference. Assets should be layered based on value and business conditions, keeping high-risk permissions and long-lived secrets on the server.
What assessments can be performed without model samples?
We can review assets, delivery chains, credential architecture, update strategies, and validation plans, but we cannot claim that model extraction prevention, runtime protection, or performance compatibility have passed.
Want to test this on your own app?
Submit the release candidate, target systems, and critical business paths for a Yudun PoC and compatibility assessment.
Continue with: Runtime security boundaries for mobile AI applications