How to Vet a Software Development Partner With a Paid Trial Sprint
Picture month three of a new vendor contract. The senior engineer from the sales call has quietly rotated off, pull requests arrive in 3,000-line batches, and nobody can explain why the tests were skipped.
Each of those problems can surface in week one. When you vet a software development partner, the most reliable evidence isn't the portfolio or the pitch deck. It's a few weeks of their engineers working in your codebase, under your rules.
Below: a setup checklist, a week-by-week plan, and a reusable scorecard.
What is a paid trial sprint, and why does it beat a portfolio review?
During a paid trial sprint, a software vendor's dedicated developers complete an actual item from your backlog inside your live codebase for a short, compensated duration. More than just completing a single feature, this engagement lets you directly evaluate their real-time workflows and commit history.
Standard portfolios merely showcase past deliverables under unverified client conditions, while customer references reflect subjective impressions. Neither reveals who actively authors the code, how pull requests are reviewed, or how the team responds when requirements pivot.
| Evidence | What the sales process gives you | What a trial sprint gives you |
|---|---|---|
| Team | Bios and résumés | Commits from the engineers you interviewed |
| Code quality | Screenshots and case studies | Pull requests you can read line by line |
| Process | "We follow agile best practices" | Review comments, CI runs, and ticket history |
| Estimates | A proposal number | Estimated vs. actual time on a real task |
What should you settle before day one?
Converting a trial into a true test relies on six core criteria. These should be the initial questions presented to any prospective software outsourcing vendor; any hesitation or pushback from them provides valuable insight before committing any funds.
- Name the engineers. Interview the specific people who'll write the code, live, before the trial starts. The trial only means something if the people you evaluate are the people you keep.
- Work in your repository. Give each engineer an individual account in your own GitHub or GitLab organization. A vendor-hosted repo that ends in a bulk upload hides what the trial should show.
- Write the definition of done. Share acceptance criteria, test expectations, and review requirements up front, at the standard you hold your own team to.
- Fix the overlap window. Agree in writing on the daily hours when both teams are online together. With engineers in the Philippines (UTC+8), overlap is a staffing decision rather than a fixed fact of geography.
- Sign a copyright assignment first. Under U.S. law (other countries differ), commissioned work counts as "made for hire" only if it fits one of nine statutory categories, and software isn't one of the listed categories (Copyright Office, Circular 30). A signed assignment covering all trial code is the safer route, and our offshore development red flags guide covers what to ask counsel for.
- Keep real user data out. Run the trial against staging with synthetic data. For EdTech teams, stripping names from student records doesn't make them de-identified under FERPA, so don't hand trial engineers production exports.
No engagement model yet? Our software outsourcing models guide compares the options.
How do you structure a 2 to 4 week paid trial sprint?
Build the trial around one real backlog item, one environment, and a weekly demo. Use two weeks for a contained task and four when the work needs integration and a second review cycle.
Pick a task that's real but not critical
Select actual, deliverable tasks that intersect with core systems—such as your authentication infrastructure, database architecture, or CI/CD workflow—preventing the partner from side-stepping technical challenges. However, isolate this scope from critical product launch dates, as evaluating a high-stakes emergency gauges crisis handling rather than standard execution.
Set up the repo so it enforces the rules
Protect your main branch before the first commit lands. On GitHub, branch protection can require approving reviews and passing status checks before any pull request merges (GitHub Docs).
- Close the admin loophole. Branch protection doesn't apply to admins by default, so give vendor engineers write access only and enable "Do not allow bypassing the above settings."
- Require signed commits. Only signed, verified commits can then reach the protected branch, which ties each one to the account whose key signed it.
- Require a second approver. GitHub can require someone other than the last pusher to approve, so no engineer signs off on their own last-minute changes.
- Pin required checks to your CI. Anyone with write access can set a status check's state, but GitHub lets you require that a check come from a specific app, such as your CI provider.
Follow a week-by-week plan
| Week | What happens | What you're watching |
|---|---|---|
| Week 1 | Access, environment setup, first small PR opened and reviewed | Onboarding speed and the quality of early questions |
| Week 2 | Core build in small PRs, first demo on staging | Commit cadence, review depth, CI discipline |
| Week 3 (4-week trials) |
Integration, added tests, second demo | Whether earlier feedback shows up in the code |
| Final week | Handover demo and a written debrief from both sides | Estimate vs. actual, candor about what didn't work |
Budget it like normal work
Pay the long-term rate, since a discount mostly tests which vendor can afford to subsidize you.
The math is engineers × hourly rate × hours: two engineers at an illustrative $40 per hour for two 40-hour weeks costs $6,400. For current ranges by seniority, see our Philippines developer rates guide.
Not sure your trial scope will tell you what you need to know?
Book a free first session with a senior Hireplicity architect and pressure-test it before you brief vendors.
Book a Free First Session →How do you vet a software development partner's work? The Repo Evidence Test
The Repo Evidence Test is a five-signal scorecard for judging a paid trial project using only artifacts in your own repository. Think of it as a technical due diligence checklist you run yourself: score each signal pass or fail.
Treat a fail on signal 1 or 4 as a walk-away, because both mean the evidence itself can't be trusted.
| Number | Signal | Where to look | Pass | Fail |
|---|---|---|---|---|
| 1 | Authorship | Commit history and contributor list | Signed commits from the engineers you interviewed | Unknown contributors, or one account committing for several people |
| 2 | Cadence | PR timeline | Small PRs opened throughout the sprint | One large PR the day before the demo |
| 3 | Review depth | PR conversations | Comments on design, edge cases, and tests, with feedback addressed | Silent approvals, or comments left unresolved |
| 4 | Pipeline discipline | Checks on each merged PR | Required checks ran and passed, and failures were fixed in code | Checks disabled, skipped, or bypassed to merge |
| 5 | Demo reality | Staging environment | Working software you can click through each week | Slides, screenshots, or "it works on my machine" |
Signals 1 and 4 are marked in red: a fail on either should end the evaluation.
Signal 1, authorship, is the bait-and-switch check. If commit authors don't match your interview list, you're evaluating the wrong team.
Signal 2, cadence, hints at whether a team plans work before writing it. Google's engineering practices call about 100 lines a reasonable change and 1,000 lines usually too large, and they let reviewers reject a change for size alone (Google eng-practices).
Signal 3, review depth, lives in PR conversations, not approval badges. Google's guidelines expect new or updated tests with any change that adds or modifies logic, so look for reviewers asking for them.
Signal 4, pipeline discipline, has a trap worth knowing.
On GitHub, a skipped job reports its status as "Success" and won't block a merge, even when it's a required check (GitHub Docs). Search the commit messages for the skip-checks: true trailer, and ask about every skipped run.
Signal 5, demo reality, is the plainest check. No working software on staging by the second weekly check-in is a gap no status deck can close.
Outside the repo, compare estimated against actual time and ask for the gap explained in writing. Then review what engineers asked before building, since those questions show how a team handles ambiguity.
DORA's research adds context: it finds speed and stability aren't tradeoffs and names smaller batch sizes as a common way to improve both (DORA). PR size and review time show those habits early.
Also ask the vendor to disclose which AI coding tools touched the trial code. Our guide to AI-generated code quality covers the contract terms and controls to require before you scale.
Which trial results should end the evaluation?
Some trial problems are coachable. The ones on the left are a decision, not a discussion.
| Walk away | Fixable with feedback |
|---|---|
| An interviewed engineer is replaced mid-trial without notice | Early PRs are too large but shrink after you ask |
| Commits come from people you never met | Test coverage is thin but improves in the next PR |
| Required checks are disabled or bypassed to get a merge through | Estimates miss, and the vendor explains why in writing |
The dividing line is response to feedback: fixable signals improve after one round of it.
For contract-stage warning signs, see our red flags guide.
What happens after the trial ends?
End every trial with one of three calls: pass, extend, or convert.
Revoke the vendor's repository and tool access the same day.
By a week or two only when the signals look good but are too thin to judge.
Negotiate the long-term contract during the trial so a "yes" doesn't stall in legal review.
How does Hireplicity run a first sprint?
We think a first sprint should look like the hundredth. Our engineers join your Slack, Jira, and GitHub in week one, follow your PR review process, and target production-ready pull requests merged in your repo within four weeks of the first call.
Our Philippine teams provide up to four hours of daily overlap as standard, or can work your shift. Taylor Basilio, our U.S.-based CEO, personally reviews every engagement in its first 30 days, and if an engineer isn't performing, we replace them within the week at our cost.
Frequently asked questions
Two to four weeks covers most vetting decisions. Two weeks suits a contained task, while four leaves room for integration work and a second review cycle. If the entire project fits in that window, skip the trial and run the project itself with milestone payments.
Yes. Paying gives you the standing to demand production-quality work, and it gives the vendor a reason to staff the trial with the engineers you'll actually keep. Pay the rate you'd pay long-term, because a discounted trial mostly tests which vendor can afford to subsidize the work, not which one fits.
Cover people, code, and contract. Confirm the named engineers, then check commit authorship, pull request size, review depth, CI enforcement, and weekly staging demos in your own repository. Pair those checks with a signed copyright assignment, named staffing terms, and exit terms, all settled before any code is written.
Yes, but give both vendors the same task, the same acceptance criteria, and the same repository rules. Otherwise you're comparing tasks rather than teams. Running two trials also doubles the review load on your engineers, so make sure someone on your side has time to read every pull request.
Ask who will write the code and whether you can interview them this week. Ask whether they'll work in your repository under your branch protection rules, and what "done" means for a pull request. Treat a vague answer to any of these as a reason to dig deeper before you sign anything.
Conclusion
Effective vetting of a software development partner requires a fundamental mindset change: focus on tangible commitments rather than sales pitches. By executing a paid trial sprint directly within your repository under your specific engineering standards, you convert verbal assurances into verifiable data.
Utilize the Repo Evidence Test to grade performance, systematically eliminate vendors exhibiting uncoachable red flags, and finalize agreements backed by concrete performance metrics rather than optimistic assumptions.
Interested in observing a trial sprint inside your codebase?
Connect with Hireplicity to review your technical requirements and timeline, and receive a curated roster of pre-vetted engineers aligned with your tech stack within five business days.
Connect with Hireplicity →- GitHub Docs, "About protected branches." — https://docs.github.com/en/repositories/configuring-branches-and-merges-in-your-repository/managing-protected-branches/about-protected-branches
- GitHub Docs, "Status checks." — https://docs.github.com/en/pull-requests/reference/status-checks
- Google Engineering Practices, "Small CLs." — https://google.github.io/eng-practices/review/developer/small-cls.html
- DORA (Google Cloud), "DORA's software delivery performance metrics," last updated January 5, 2026. — https://dora.dev/guides/dora-metrics/
- U.S. Copyright Office, "Works Made for Hire" (Circular 30). — https://www.copyright.gov/circs/circ30.pdf

