Your AI-built web3 app before mainnet: what to check first
By Oleksii Skurikhin · Delivery Manager
- Published
- 10 min read

You built a blockchain app with AI tools and want to launch it. The order we suggest for checking it: whether you need a blockchain, the contract, the keys, then a testnet and an audit, with a sheet to fill in.
Key takeaways
- A 3 February 2026 arXiv study ran one static analysis tool over 1,033 AI-written contracts from three agents: 47.4% to over 75% of an agent's contracts had at least one finding.
- First ask whether the app needs a blockchain: NIST's 2018 overview lists the features that suit one and further factors to weigh, such as public data and no removal of the original data.
- Keep private keys out of the repository, prompts and chats: NIST says a transfer made with a stolen key generally cannot be undone. No single account should control the contract.
- Before mainnet a contract needs tests, analysis tools, a testnet rehearsal and an independent audit. ethereum.org says audits will not catch every bug and deployed code usually cannot be patched.
- Work on a copy, on a testnet, with throwaway keys, for 60 minutes. The sheet at the end records what you still do not know.
On 3 February 2026, researchers at Deakin University posted a study of Solidity smart contracts that three AI coding agents had written. The authors report that the contracts were syntactically correct and functionally complete, yet many carried security flaws that could be exploited in real-world settings. What they counted were findings of one static analysis tool, not attacks (arXiv study of LLM-generated smart contracts, 3 February 2026).
If you built a blockchain app with AI tools and it works on a test network, the next stretch runs to real money. Mainnet is the real network, where transactions move real value. We suggest this order: whether you need a blockchain at all, what the contract lets people do, where the keys live, and what has to happen on a testnet and in an audit before mainnet.
Where do I start without risking real money?
Work on a copy of the project, made without .env files or any file that holds a real key, on a test network only, with new accounts that hold nothing of value. Do not deploy to mainnet, do not move real funds, and do not paste a real key anywhere while you check. Stop after 60 minutes and write down what you still do not know.
The ethereum.org Networks page says it is not recommended to reuse mainnet accounts on testnets, or the other way round, so make fresh accounts for this check.
Does my app need a blockchain at all?
A blockchain is possibly not needed, and that question comes before any Solidity. NIST's 2018 overview says many organisations start from a wish to use a blockchain somewhere and then look for a place. It recommends the reverse: understand where the technology fits, then find the systems that match (NIST IR 8202, Blockchain Technology Overview, October 2018).
NIST lists features that make a blockchain suitable, among them many participants, distributed participants, and a want or need for no trusted third party. It also lists further factors to weigh: data on a permissionless network is generally public, there is no removal process for the original data, and many permissionless networks process transactions more slowly than other IT systems.
Answer three questions in writing. They are our advice, not NIST's rule.
- Who writes to the record, and do the readers trust that writer? If it is only your app, and your users trust you, the features NIST lists are not what the app needs.
- Does anyone outside your company have to verify the record without asking you?
- Would you be comfortable with this data being public and permanent?
If the answers are "only us", "no" and "no", compare a regular database with backups before you spend another week on contracts.
What did the Deakin study find in AI-written Solidity?
In the Deakin study the authors report 1,033 contracts from three coding agents, checked with Slither, a static analysis tool. Between 47.4% and over 75% of an agent's contracts had at least one vulnerability, depending on the agent, and low-severity findings were the largest group for all three agents (arXiv study).
The 34 prompts were worded the way an inexperienced developer might describe an idea, and each was run ten times per agent. The paper's own count does not add up to exactly that design, so treat the total as approximate.
Those are findings of one tool on single contracts, so read them as a pattern for your app, not a rate. The categories counted include lack of input validation, denial of service, insecure randomness, arithmetic issues, unchecked external calls, reentrancy, logical errors and access control. The authors also report that more lines of code meant more findings.
The OWASP Smart Contract Top 10 for 2026 is a general list of smart contract risks, not a list about AI-written code. It ranks access control vulnerabilities first and business logic vulnerabilities second, and it builds on 2025 incident data covering 122 smart contract incidents.
Turn these into five questions for the person reviewing your contract.
- Who can call each function that moves money or changes who is in charge? ethereum.org notes that any account can call a public or external function, which is a problem if anyone can then perform a sensitive operation such as minting tokens (ethereum.org, Smart contract security).
- Does the money math follow rules you can state in plain sentences? OWASP describes business logic flaws as design-level flaws that break the intended rules. Write each rule as a sentence and ask for one test per rule.
- Is a balance updated before money is sent out? ethereum.org explains that a malicious contract can call back into a vulnerable one before the first call finishes, and recommends the checks-effects-interactions order: checks, then state changes, then calls to other accounts.
- Which Solidity version does the code use? Since version 0.8.0 the compiler rejects code that overflows or underflows an integer; older versions need explicit checks or a library (ethereum.org). The line that starts with
pragma solidityshows which versions a file accepts; the version actually used is set in the project's build settings, so ask for both. - How much of this code is needed? The study says AI models sometimes add unnecessary code and sometimes invent functions that do not exist. ethereum.org advises reusing audited libraries such as OpenZeppelin Contracts.
To get a first inventory, paste this into your AI tool, working in a copy of the project.
Read only. Do not edit, create or delete any file, and do not run any command.
Do not open .env files and do not print any key, seed phrase or secret: give the file name and line number only.
Read the Solidity files in CONTRACTS-FOLDER and write a table with one row per public or external function: its name, who is allowed to call it, whether it moves ETH or tokens, whether it changes an owner, admin or upgrade address, and which external calls it makes.
Then list: the Solidity version range on the pragma line of each file; every place a state variable is updated after an external call; every owner, admin or role address and where it is set; every place the code reads a private key, a seed phrase or an environment variable, with file name and line number only.
Do not suggest fixes. If something cannot be answered from the code, write UNKNOWN. Stop after the table and the lists.The study's authors write that an LLM-based self review by developers cannot replace rigorous verification and independent auditing.
Where do the keys and wallets live?
A private key is what proves that you may move funds or change a contract, so it never goes into the repository, a prompt, a chat or a rules file for an AI tool. NIST says that whoever steals a private key gets full access to all digital assets that key controls, and that a transfer made with it generally cannot be undone.
OWASP's list of attack vectors beyond contract code says that many of the largest losses of 2025 came from off-chain and operational threats, among them multisig manipulation, supply chain attacks, drainer malware and phishing.
Four rules follow.
- Use separate keys for separate jobs: a throwaway key for the testnet, one for deploying, one for administering the contract. ethereum.org advises against reusing accounts between testnets and mainnet.
- Do not let one account control the contract. ethereum.org says a single owner address is a single point of failure, and that if its keys are compromised the attacker can attack the contract. It suggests multiple administrative accounts or a multisig wallet that needs a minimum number of signatures, for example 3 of 5. A multisig only helps if each signer is a different person or device, with keys guarded like any other key.
- Keep signing keys on a hardware wallet or another store that your AI tool cannot read. NIST notes that many users store private keys on special secure hardware.
- If a real key was ever in the repository, in its history, or pasted into a chat, treat it as exposed and mark the check sheet at the end NOT READY. Do not move funds during this check. Replacing the key means a new key, moving what the old one holds and removing its roles on the contract, as a separate, planned step.
To look for secrets in a copy of the project, run these two commands in your own terminal window, not through an AI tool.
grep -rIlE -i "private_key|mnemonic|seed phrase|secret" . --exclude-dir=node_modules --exclude-dir=.git
git log --all --full-history --oneline -- '*.env*'The first command lists the names of files that mention those words, and the second lists commits that touched an env file. Both can list harmless files, and an empty result means these commands found nothing, not that nothing exists.
What must happen on a testnet and in an audit before mainnet?
Our suggested order is to rehearse the contract on a testnet, have someone who did not write it review it, fix what they find, rehearse the final version again, and only then deploy to mainnet. ethereum.org says to test any contract code on a testnet before deploying to Mainnet, and that deployed contract code usually cannot be changed to patch security flaws, so every gate below happens before launch.
| Gate | What it can show | What it cannot |
|---|---|---|
| Unit tests | Functions behave as expected for the data in the tests | Cases nobody wrote a test for (ethereum.org: tests are only as effective as the tests written) |
| Static analysis and fuzzing | Known patterns flagged by tools such as Slither, Aderyn and Mythril; random inputs that break a rule | Complex and market-based vulnerabilities, which the study says Slither cannot detect |
| Testnet rehearsal | The contract and your deploy steps in a production-like environment | Behaviour with real value, because testnet ETH is supposed to have none |
| Independent audit | A second review by people who did not write the code | Every bug (ethereum.org: audits will not catch every bug) |
ethereum.org describes an emergency stop, a function that blocks calls to vulnerable functions, and suggests decentralising its control with a timelock or a multisig, because it increases the need for users to trust whoever can activate it.
Decide in writing who can pause, who can upgrade and what each of them can do. A right to pause or upgrade is also a power someone can abuse or steal, so protect it like the admin key. An upgrade path is itself something to audit: OWASP ranks proxy and upgradeability vulnerabilities tenth in its 2026 list.
Here is how the checks fit together. Picture a founder who builds a fan-rewards app with an AI tool. Members earn tokens for attending matches, and an admin function mints more.
The founder asks who can call that function, whether anyone but the allowed caller can mint and who can pause it, and writes the answers on the check sheet at the end of the page. Say the AI also puts the deployer key in a config file. That key counts as exposed, the sheet says NOT READY and a replacement is planned as a separate step. The testnet rehearsal of the final version has to show that only the allowed caller can mint, that the pause works and that the deploy steps work. A line that stays UNKNOWN means the contract waits.
What do we know and what have we not verified?
The Deakin paper's arXiv page does not say that it was peer reviewed. It used one static analysis tool on single contracts. Each contract came from one request, the prompts were themselves written by an AI model, and the paper reports no comparison with human-written contracts. Its authors say they used Slither's own classification, which leaves out flash loan vulnerabilities, and that Slither cannot detect complex and market-based vulnerabilities.
NIST's report is from October 2018 and describes blockchain in general, not smart contract coding. The study and the ethereum.org pages concern Solidity on Ethereum, while NIST and OWASP are broader. Other chains, token law, and the code of a website that connects to a wallet are not covered.
The prompt, the commands and the 60-minute time-box have not been tried on a real AI-built project, and none of this replaces an audit. If your answers to the three questions say you do not need a blockchain, skip the contract sections.
What can I check today?
Give it 60 minutes on a copy. Fill in the sheet below and mark each line OK, UNKNOWN or NOT READY; a line you could not check stays UNKNOWN.
WEB3 LAUNCH CHECK SHEET (work on a copy, testnet only, throwaway keys, stop after 60 minutes)
PROJECT:
CHAIN AND TESTNET:
CHECKED BY AND DATE:
1 NEED FOR A BLOCKCHAIN: WHO WRITES: ____ | WHO MUST VERIFY WITHOUT ASKING US: ____ | DATA MAY BE PUBLIC AND PERMANENT: YES / NO | DECISION: NEEDED / NOT NEEDED / UNKNOWN
2 CALLERS: FUNCTIONS THAT MOVE MONEY OR CHANGE CONTROL: ____ | EACH HAS A NAMED ALLOWED CALLER: YES / NO | STATUS: OK / UNKNOWN / NOT READY
3 MONEY RULES: RULES WRITTEN AS SENTENCES: ____ | TEST PER RULE: YES / NO | STATUS:
4 EXTERNAL CALLS: BALANCE UPDATED BEFORE SENDING: YES / NO / UNKNOWN | STATUS:
5 SOLIDITY VERSION: ____ | CODE I CANNOT EXPLAIN: ____ | STATUS:
6 KEYS: KEY WORDS FOUND IN CODE OR HISTORY: ____ | KEYS PER JOB: TESTNET / DEPLOY / ADMIN | ADMIN IS A MULTISIG: YES / NO | STATUS:
7 TESTS AND TOOLS: UNIT TESTS RUN: YES / NO | STATIC ANALYSIS RUN: YES / NO | TESTNET REHEARSAL OF THE DEPLOY STEPS: YES / NO | STATUS:
8 AUDIT AND FAILURE PLAN: INDEPENDENT AUDIT DONE: YES / NO | WHO CAN PAUSE: ____ | WHO CAN UPGRADE: ____ | MONITORING: YES / NO | STATUS:
NOT OK COUNT: ____
NOTES: EVERY UNKNOWN AND NOT READY LINE, WITH WHAT YOU SAWNext step
Sources
Checked on 11 October 2026.
About the author

Category:Software DevelopmentReview and trust in AI code
Back to Blog

