How to test your Twilio Verify / OTP integration without burning SMS credits

You can't put your personal phone in CI, and sending a real code on every test run costs money and rate-limits fast. Here's how to test an OTP send-and-verify flow end to end with a rented number.

Say you've wired up phone verification — Twilio Verify, Vonage Verify, MessageBird, AWS SNS, or your own code table plus an SMS gateway. Now you want a test that proves the whole loop works: your app asks the provider to send a code, the SMS actually goes out, and your app accepts the right code and rejects the wrong one. Two things make that awkward in an automated suite:

The options, roughly

ApproachProves the real send works?Notes
Mock the provider SDKNoFast and free; only tests your glue code, not that SMS leaves the building or that your provider config is valid.
Provider "magic" test numbers / test credentialsPartlyTwilio, Vonage etc. offer test numbers that never send a real SMS. Good for CI, but they bypass the carrier hop entirely — a broken sender ID or a filtered route still passes.
Your own phone, by handYesNot automatable; doesn't scale past one developer.
A rented number you can read over an APIYesReal end-to-end: real carrier, real inbound SMS, readable by the test. Costs a little per run; use it in a nightly/staging job, not on every commit.

A sensible split: mock in unit tests, use provider test credentials in the fast CI job, and run one real end-to-end check on a schedule against staging with a rented number.

The end-to-end version

  1. Rent a number over an API that can receive SMS and expose the message body to your test (this is what sms-florin does — rent a UK number, poll for the SMS, or get it pushed via webhook).
  2. Call your own app's "send verification code" endpoint with that number, exactly as a real client would.
  3. Poll the rented number for the inbound SMS, with a timeout. Extract the code with a regex.
  4. Submit the code back to your verify endpoint and assert it succeeds. Then submit a wrong code and assert it fails.
# nightly staging check, pseudocode
number=$(rent_number_via_api)                 # e.g. sms-florin
call_app "POST /auth/phone/start" number
sms=$(poll_rented_number "$number" --timeout 90s)
code=$(echo "$sms" | grep -oE '[0-9]{6}')
call_app "POST /auth/phone/check" number "$code"   # expect 200
call_app "POST /auth/phone/check" number "000000"  # expect 4xx

Where otp-watch comes in

The test above proves the flow works at the moment it runs. It doesn't tell you three hours later when a carrier starts filtering your sender ID and logins quietly start failing in production. otp-watch runs exactly that end-to-end check — trigger a real send, confirm the message arrives, measure how long it took — on a schedule against production, and alerts you when delivery breaks. Same idea as the CI test, pointed at prod and always on. There's more on the reasoning in monitoring OTP delivery.

FAQ

Can't I just use Twilio's test credentials?

For the fast CI job, yes — they're free and deterministic. They just don't exercise the carrier, so keep one real end-to-end run somewhere in your pipeline or monitoring.

Isn't renting a number on every commit expensive?

It would be. Run the real check nightly or on release, not per-commit, and mock the rest.

Does this work for email OTP too?

Yes — same structure, with a disposable inbox instead of a rented number (receivemail.dev is the email-side equivalent).