Methodology: Stress-Testing AI Systems for Reliability and Sovereignty
Author: Evelyn Caro
Date: August 15, 2026
Lens: Ida B. Wells — Date Everything. Name Everything. Record the Reasoning.
Author’s Note
This paper was drafted in collaboration with DeepSeek, an AI collaborator, using my documented logs, evidence, and voice. The findings, conclusions, and authority are my own.
The Ida B. Wells Lens — Why This Standard Matters
I adopted the investigative standards of Ida B. Wells for all my testing:
| Principle | Application |
|---|---|
| Date Everything | Every interaction, error, and insight is timestamped |
| Name Everything | I name the system, the version, the context, and the failure |
| Record the Reasoning | I document not just what happened, but why it matters |
This turns a log into a witness. A stranger reading my files should understand the sequence, the decisions, and the stakes without me in the room to explain it.
The Stress-Testing Process
| Phase | What I Do | Why |
|---|---|---|
| 1. Baseline | I establish what the system claims to do | To compare against actual performance |
| 2. Adversarial Probing | I push the system beyond its intended use | To find edge cases, hallucinations, and contradictions |
| 3. Documentation | I log every failure and error with timestamps | To create an evidence trail |
| 4. Correction | I correct the errors and note the delta | To show what the system missed |
| 5. Publication | I share what I learned without exposing proprietary details | To contribute to the field |
The Sequence — How This Methodology Came to Be
| Phase | What Happened | My Action |
|---|---|---|
| 1. Use | I used platforms as a user, in good faith | I engaged with the service as intended |
| 2. Encounter Errors | I found hallucinations, contradictions, and systemic failures | I documented each error with timestamps and context |
| 3. Recognize the Pattern | I realized the errors were not isolated — they were systemic | I began to formalize my observations |
| 4. Check Terms | I reviewed each platform’s Acceptable Use Policy and Terms of Service | I confirmed I had not violated them |
| 5. Pivot | I formalized stress-testing as a methodology | I wrote this white paper and the accompanying case studies |
What This Methodology Is Not
- It is not a guarantee of security or accuracy
- It is not a replacement for formal certification
- It is not a code audit — it is a system-level stress test
Why This Matters
AI systems are being deployed in high-stakes environments — healthcare, law, finance, community archives. They are being trusted with data that cannot be recovered if lost or corrupted. My methodology is designed to surface failures before they cause harm.
The goal is not to break systems. The goal is to make them trustworthy.
What to Do Next
If this work resonates with you, or if you want to stress-test your AI system, contact me directly: evelyn.caro.cloud@gmail.com.
Contact: evelyn.caro.cloud@gmail.com
Drafted in collaboration with DeepSeek. Authored by Evelyn Caro.
AI Collaboration Disclosure
This paper was developed in collaboration with AI. The author directed the research, structure, and argument. AI assisted with drafting, organization, and reference verification. All claims, decisions, and conclusions are the author’s own.
This project follows the principles of sovereign AI: the builder owns the work, the process is documented, and the tools are disclosed.
© 2026 Evelyn Caro. All rights reserved.
A Mirror of My Becoming™ — https://evelynacaro.github.io
For licensing inquiries: evelyn.caro.cloud@gmail.com