Verifier mutation scores¶
For each sample task, one plausible bug at a time was planted in the reference solution (operator
and comparison flips, constant tweaks, dropped statements, negated conditions, ...) and graded
against the task's cases. A killed mutant fails at least one case; a survivor passes them all
(an equivalent program, or a gap in the cases). Mutants that do not compile or do not change the
program are invalid. Generated by tools/mutation_table.py.
| Task | Mutants | Killed | Survived | Invalid | Score |
|---|---|---|---|---|---|
py-balanced-brackets |
24 | 24 | 0 | 0 | 100.0% |
py-flatten |
14 | 14 | 0 | 0 | 100.0% |
py-gcd-lcm |
11 | 11 | 0 | 0 | 100.0% |
py-matrix-rotate |
12 | 12 | 0 | 0 | 100.0% |
py-merge-intervals |
14 | 14 | 0 | 0 | 100.0% |
py-parse-duration |
21 | 18 | 1 | 2 | 94.7% |
py-rle-encode |
17 | 17 | 0 | 0 | 100.0% |
py-roman |
40 | 40 | 0 | 0 | 100.0% |
py-semver-compare |
35 | 30 | 2 | 3 | 93.8% |
py-slugify |
11 | 11 | 0 | 0 | 100.0% |
py-two-sum |
6 | 4 | 0 | 2 | 100.0% |
py-valid-ipv4 |
30 | 30 | 0 | 0 | 100.0% |
py-word-frequency |
4 | 4 | 0 | 0 | 100.0% |
| All tasks | 239 | 229 | 3 | 7 | 98.7% |
Survivors (3)¶
Each survivor is either behaviourally equivalent to the original or points at a case the task does not have. They are listed so that a human can decide which.
py-parse-duration, constant bool (line 13)
@@ -11 +11 @@
- return sum((int(g) * s for g, s in zip(match.groups(), _SECONDS, strict=True) if g is not None))
+ return sum((int(g) * s for g, s in zip(match.groups(), _SECONDS, strict=False) if g is not None))
py-semver-compare, comparison flip (line 26)
@@ -23 +23 @@
- return -1 if pre_a < pre_b else 1
+ return -1 if pre_a <= pre_b else 1
py-semver-compare, comparison flip (line 19)
@@ -16 +16 @@
- return -1 if nums_a < nums_b else 1
+ return -1 if nums_a <= nums_b else 1