Say you've wired up phone verification — Twilio Verify, Vonage Verify, MessageBird, AWS SNS, or your own code table plus an SMS gateway. Now you want a test that proves the whole loop works: your app asks the provider to send a code, the SMS actually goes out, and your app accepts the right code and rejects the wrong one. Two things make that awkward in an automated suite:
| Approach | Proves the real send works? | Notes |
|---|---|---|
| Mock the provider SDK | No | Fast and free; only tests your glue code, not that SMS leaves the building or that your provider config is valid. |
| Provider "magic" test numbers / test credentials | Partly | Twilio, Vonage etc. offer test numbers that never send a real SMS. Good for CI, but they bypass the carrier hop entirely — a broken sender ID or a filtered route still passes. |
| Your own phone, by hand | Yes | Not automatable; doesn't scale past one developer. |
| A rented number you can read over an API | Yes | Real end-to-end: real carrier, real inbound SMS, readable by the test. Costs a little per run; use it in a nightly/staging job, not on every commit. |
A sensible split: mock in unit tests, use provider test credentials in the fast CI job, and run one real end-to-end check on a schedule against staging with a rented number.
# nightly staging check, pseudocode number=$(rent_number_via_api) # e.g. sms-florin call_app "POST /auth/phone/start" number sms=$(poll_rented_number "$number" --timeout 90s) code=$(echo "$sms" | grep -oE '[0-9]{6}') call_app "POST /auth/phone/check" number "$code" # expect 200 call_app "POST /auth/phone/check" number "000000" # expect 4xx
The test above proves the flow works at the moment it runs. It doesn't tell you three hours later when a carrier starts filtering your sender ID and logins quietly start failing in production. otp-watch runs exactly that end-to-end check — trigger a real send, confirm the message arrives, measure how long it took — on a schedule against production, and alerts you when delivery breaks. Same idea as the CI test, pointed at prod and always on. There's more on the reasoning in monitoring OTP delivery.
For the fast CI job, yes — they're free and deterministic. They just don't exercise the carrier, so keep one real end-to-end run somewhere in your pipeline or monitoring.
It would be. Run the real check nightly or on release, not per-commit, and mock the rest.
Yes — same structure, with a disposable inbox instead of a rented number (receivemail.dev is the email-side equivalent).