מפתח עמודים ב-OpenStax
איזה עמוד בספר Introductory Statistics 2e שווה לקחת ממנו חומר, ולאיזה נושא באפליקציה
מספרי העמודים כאן הם עמודי ה-PDF (1–847). מספר העמוד המודפס בספר עצמו נמוך ב-12: עמוד PDF 348 הוא עמוד 336 בספר. הקובץ:
~/Downloads/introductory-statistics-2e_-_WEB.pdfהספר הוא CC BY 4.0, ולכן מותר לעבד ממנו תרגילים ודוגמאות.
המסמך נבנה משלוש סריקות נפרדות של הספר, פרק אחר פרק, מול התוכן שכבר קיים באפליקציה. כל המלצה מציינת מה יש שם ומה זה מוסיף — עמוד שרק חוזר על מה שכבר כתוב לא נכנס לרשימה. בכל דוח יש גם קטע "SKIPPED" של מה שנבדק ונדחה, והוא שימושי לא פחות: הוא חוסך חזרה לאותם עמודים.
ארבע נקודות שעולות מכל שלוש הסריקות
- ליחידות 1–4 של המפרט (מושגי יסוד, הצגת נתונים, מדדי מרכז, מדדי פיזור) נבנה עכשיו נושא "יסודות", ופרק 2 של הספר הוא ההתאמה הישירה ביותר לו — הפרק העשיר ביותר בספר כולו לצרכינו.
- את פרק 10 אסור למחזר מבחינת נוסחאות. הספר מלמד את גרסת Aspin-Welch הלא-משוקללת כברירת מחדל ואף כותב במפורש "אין לאחד את השונויות", בעוד שהבחינה דורשת את מבחן המשוקלל עם . לקחת משם תרחישים בלבד.
- מבחני פרופורציות אינם על הבחינה — אין להם יחידה במפרט, ואינם מופיעים בקטלוג המבחנים ביחידה 38. כל הסעיפים עליהם נדחו.
- החוסר האמיתי הגדול ביותר הוא מבחן חי-בריבוע לאי-תלות עם שכיחויות צפויות שאינן שוות, ו-ANOVA עם קבוצות בגדלים שונים — שני מקרים שהאפליקציה כרגע אינה מכסה כלל.
OpenStax Introductory Statistics 2e — mining report, chapters 1, 2, 3(partial), 6, 7
Convention: PDF page first, printed book page in parentheses. PDF = printed + 12.
Headline finding
The app has no topic at all for spec units 1–4 (סולמות מדידה · הצגת נתונים · מדדי מרכז · מדדי פיזור). Chapter 2 is the single richest vein in this whole PDF and it feeds exactly that gap: mean-vs-median resistance, IQR and the 1.5×IQR outlier rule, percentiles from a cumulative-frequency table, the SS/deviation table, weighted means from a frequency table, and skew↔mean/median with real computed numbers. Chapter 1's exercise sets (not its prose) are a large ready-made bank for research-critique. Chapter 6 is nearly worthless — see SKIPPED. Chapter 7's prose is calculator-bound, but a handful of its conceptual items (individual-vs-mean) are the best material in the book for the CLT unit.
SKIPPED — looked at, deliberately rejected
| What | PDF pages | Why rejected |
|---|---|---|
| All of §6.1 empirical rule (Ex 6.3, 6.5, 6.6) | p350–p352 (338–340) | Duplicates z-scores/lesson.md line-for-line, including the 68-95-99.7 figure. The app's version (with the six-segment breakdown) is better. |
| All of §6.2 "Using the Normal Distribution" (Ex 6.7–6.12) | p352–p358 (340–346) | Every answer comes from normalcdf/invNorm. מתא"ם gives no calculator and no Z table. Unusable. |
| §7.2 CLT for Sums entirely | p382–p385 (370–373) | ΣX ~ N(nμ, √n σ) is not on this exam. |
| Normal approximation to the binomial + continuity correction | p392–p393 (380–381) | Out of scope. |
| §1.3 levels-of-measurement prose | p38–p39 (26–27) | Mediocre and partly wrong ("ordinal scale data cannot be used in calculations"). The spec's own trap list (IQ/Celsius/year-of-birth, 1=גבר 2=אישה) is sharper. Only the exercise bank (below) is worth taking. |
| §2.1 stem-and-leaf + §2.2 histogram construction | p78–p97 (66–85) | Bin-boundary arithmetic (why 59.95, how to pick bar width) is a construction drill. Exam asks you to read a histogram, never to build one. |
| The two percentile formulas i = (k/100)(n+1) and (x+0.5y)/n | p102, p104 (90, 92) | Arbitrary interpolation conventions that differ between textbooks. Teaching them would create wrong answers on real exam items. Keep the concept (Ex 2.15/2.17 below), drop the formulas. |
| §2.4 box-plot construction steps | p106–p110 (94–98) | Construction, not reading — except Figure 2.14, kept below as a figure idea. |
| Ch 3 §3.1, §3.3, §3.5 (tree/Venn), Muddy Mouse Ex 3.22 | p180–p208 (168–196) | The app has no probability topic and the spec rates these Tier B/C with an explicit "אל תפתח יותר". Only Ex 3.21 (below) earns a place, and only because it feeds χ² expected frequencies. |
| §1.4 Ethics / Stapel fraud / IRB / informed consent | p48–p49, p62 (36–37, 50) | Genuinely interesting (Stapel is a psychology fraud case, retracted 55 papers) but there is no ethics unit in the 44-unit spec. Noted at the bottom as an optional aside only. |
| All Stats Labs and Collaborative Exercises (§1.5, §1.6, §2.8, §6.3, §6.4, §7.4, §7.5) | various | Classroom activities, per brief. |
| Chapter Reviews / Key Terms / Formula Reviews / Bringing It Together | p147–p152, p166+ | Restatement. |
| All TI-83/84 boxes | throughout | Per brief. |
| Ex 2.35 GPA comparison (John 2.85 vs Ali 77) | p129 (117) | The app's z-same-score-two-classes.svg example already does this better. The only new wrinkle (comparing two negative z-scores) is small; picked up cheaply via Ex 71 instead. |
| Ex 7.92 water-taxi weight limit | p409 (397) | Excellent scenario, but it is a CLT-for-sums problem. Out of scope. |
NEW TOPIC — descriptive statistics (spec units 1–4; app has nothing)
p111–p112 (printed 99–100) — Ex 2.27, and Ex 2.26 (ER patient ages)
What is there: A town of 50 people: one earns $5,000,000, the other 49 earn $30,000.
Mean = $129,400, median = $30,000. Directly above it, Ex 2.26 gives 40 sorted ER-patient
ages with mean 23.6 and median 24 — a case where the two agree.
Why it improves the app: The app asserts "הממוצע רגיש לחריגים, החציון עמיד" in prose
(z-scores/lesson.md) but never demonstrates it numerically. This is the cleanest possible
demonstration — the mean lands at a value no one in the town earns, and it is verifiable
in your head. The Ex 2.26 contrast (mean ≈ median in a near-symmetric set) gives the paired
control case.
Use for: lesson prose | worked problem | quiz questions (~4)
p112–p113 (printed 100–101) — Ex 2.29 and TRY IT 2.29, "which measure of center?"
What is there: (a) A weight-loss program advertises a mean loss of 6 lb in week one;
the mode is 2 lb. (b) Factory annual earnings: mode = $25,000 occurring 150 times out
of 301, median = $50,000, mean = $47,500. "What would be the best measure of the center?"
Why it improves the app: (b) is a genuine three-way conflict where the mode is held by an
absolute majority of the workforce yet sits $25,000 below both other measures — so the
"correct" answer is arguable and the reasoning is what matters. That is precisely the מתא"ם
"מה מותר להסיק" shape, and it doubles as a reading-results item on a misleading headline
statistic. The app has no "which measure of center" material at all.
Use for: quiz questions (~3) | worked problem | reading-results problem bank
p99 (printed 87) — Ex 2.13, IQR and the 1.5×IQR outlier rule
What is there: 13 real-estate prices. Ordered, M = 488,800, Q1 = 308,750, Q3 = 649,000, IQR = 340,250, fences at −201,625 and 1,159,375, so $5,500,000 is flagged as an outlier. Every number is shown. Why it improves the app: The app has no IQR anywhere — not in the lesson, not in the transformation matrix (which does list תחום בין-רבעוני but never defines it), not in a figure. Spec unit 4 subtopic 1 requires it. This gives a fully-worked instance with an outlier that is obviously an outlier, so the rule is learned rather than memorised. Use for: lesson prose | worked problem | quiz questions (~3)
p100 (printed 88) — Ex 2.14, day class vs night class test scores
What is there: Two full data sets (n = 22 each). Five-number summaries: day 32 / 56 / 74.5 / 82.5 / 99 (IQR 26.5); night 25.5 / 78 / 81 / 89 / 98 (IQR 11). Applying the 1.5×IQR rule: the day class, with more than double the IQR, has no outliers; the night class, far more tightly packed, has two (45 and 25.5). Why it improves the app: This is a built-in trap of exactly the type the spec asks for ("שני מדגמים עם אותו טווח ושונות שונה מאוד", unit 4). The larger-spread group having fewer outliers is counter-intuitive and forces the student to see that "outlier" is defined relative to the group's own spread — the same relativity idea that underlies z-scores. It also supplies the data for the comparative box-plot figure below. Use for: worked problem | quiz questions (~3) | figure
p143–p146 (printed 131–134) — §2.6 Practice, exercises 49–68
What is there: ~20 short items on skew ↔ mean/median/mode. Items 49–51 are three small data lists to classify. Items 55–63 are nine small histograms (Figures 2.32–2.40) asking "describe the shape" / "describe the relationship between the mode and the median". Items 64–68 are conceptual:
- 64: for 3;4;5;5;6;6;6;6;6;7;7;7;7;7;7;7 the mean equals the median — is the data perfectly symmetrical? Why or why not? (No.)
- 65 / 66: which of mean, mode, median is greatest / least, for a given list.
- 67: which of the three reflects skewing the most, and why.
- 68: in a perfectly symmetrical distribution, when would the mode be different from the
mean and median? (When it is bimodal.)
Why it improves the app: Items 64 and 68 correct the app's own table.
z-scores/lesson.mdstates flatly "סימטרית → שלושתם מתלכדים", which is false for a symmetric bimodal distribution, and it never warns that mean = median does not imply symmetry. Both are classic "מה נובע / מה לא נובע" traps. The nine histograms are also the cheapest possible source of visual classification drills. Use for: quiz questions (~12) | lesson correction | figure ideas (small histograms)
p115–p116 (printed 103–104) — Ex 2.31, Terry / Davis / Maris dot plots
What is there: Three 10-value letter-count samples plotted as dot plots (Figures 2.19–2.21):
Terry right-skewed (mean 3.7, median 3), Davis left-skewed (mean 2.7, median 3), Maris
symmetric (mean 4.6, median 4). The stated conclusion: the median is always closest to the
high point (the mode), while the mean is pulled out toward the tail.
Why it improves the app: The app's skew-mean-median-mode.svg is an idealised smooth
curve. This is the same lesson with 10 numbers you can add up by hand, so the student
proves the rule to themselves instead of accepting it. The raw data are small enough to
regenerate in scripts/figures.mjs and have the mean/median computed rather than drawn.
Use for: new figure (dot plots with computed M and Mdn markers) | worked problem
p121–p123 (printed 109–111) — Ex 2.32, the full deviation / SS table
What is there: Ages of n = 20 fifth-graders with frequencies. A five-column table:
value, f, (x − x̄), (x − x̄)², f(x − x̄)². Column total 9.7375; s² = 9.7375/19 = 0.5125;
s = 0.7159 ≈ 0.72. Followed on p123 by the explanation: "If you add the deviations, the sum
is always zero" — hence the squaring — and the n−1 justification.
Why it improves the app: Spec unit 4 subtopics 2–5 (סטייה מהממוצע · מדוע סכום הסטיות
תמיד אפס · סכום ריבועים SS · ÷N מול ÷(n−1)) are entirely absent from the app, and SS is the
concept that returns in ANOVA (anova-oneway uses SS without ever building it). This is a
complete, small, hand-checkable instance of the whole chain.
Use for: lesson prose | worked problem | ties anova-oneway back to a concrete SS
p114–p115 (printed 102–103) — Ex 2.30, mean of a grouped frequency table
What is there: Mean of a frequency table = Σfm / Σf where m is the interval midpoint. Professor Blount's test: 8 grade intervals, Σfm = 1,460.25, Σf = 19, mean = 76.86. A companion at p126 (printed 114), Ex 2.34 does the same for the standard deviation with a full seven-column table (Σf = 26, Σfm = 197, x̄ = 7.58, s = 3.50). Why it improves the app: Spec unit 3 asks for "5 שאלות ממוצע משוקלל" and for ΣX = M × N as a central exam manoeuvre. The app has neither. Σfm/Σf is the weighted mean, and reading it off a table is exactly the מתא"ם format. Also directly supports the "ממוצע כולל של שתי קבוצות בגדלים שונים" trap the spec flags. Use for: lesson prose | quiz questions (~4)
p101–p102 (printed 89–90) and p103 (printed 91) — Ex 2.15 and Ex 2.17, percentiles from a cumulative table
What is there: 50 statistics students' hours of sleep as a frequency / relative frequency / cumulative relative frequency table. The 28th percentile is read straight off the 0.28 entry; the median off the 0.52; Q3 = the 75th percentile is argued from "52% at or below 7, 80% at or below 8, therefore Q3 = 8". Ex 2.17 repeats it for the 80th and 90th percentiles and Q1. Why it improves the app: Cumulative relative frequency is the single most useful table skill on this exam and the app has it nowhere. It converts a percentile question into a scan of one column with no arithmetic — exactly the no-calculator technique the exam rewards. (Take the reasoning, not the index formula — see SKIPPED.) Use for: lesson prose | worked problem | quiz questions (~4)
p106 (printed 94) — Ex 2.22, Sharpe Middle School exercise minutes
What is there: 15 students report daily exercise minutes; one reports 300. Five-number summary 0 / 20 / 40 / 60 / 300. The 300 is flagged by the IQR rule (fence at 120). Removing it changes Max from 300 to 120 and leaves Min, Q1, Median, Q3 untouched. The passage then asks whether the principal is justified in buying equipment, and closes with "15 students is a small sample." Why it improves the app: A rare example that is simultaneously (a) an outlier drill, (b) a demonstration that the quartiles are resistant while the range is not, and (c) a "תוצאה מול מסקנה" decision-in-context item with a sample-size caveat. Three app topics in one 15-number data set. Use for: worked problem | quiz questions (~3) | reading-results problem bank
p108–p110 (printed 96–98) and p106 (printed 94) — Ex 2.24 / Figure 2.14; Ex 2.25; Figure 2.13
What is there: Three box-plot ideas worth generating:
- Figure 2.14 — the day and night class box plots drawn on one shared number line, making the "wider box, shorter whiskers" contrast visible at a glance.
- Ex 2.25 / Figure 2.15 — 10;10;10;15;35;75;90;95;100;175;420;490;515;515;790 → Min 10, Q1 15, Med 95, Q3 490, Max 790. A violently right-skewed box plot where the box sits far left with an enormous right whisker.
- Figure 2.13 — a degenerate box plot where Min = Q1 = 1 and Median = Q3 = 5, read as
"at least 25% of the values equal 1; at least 25% equal 5".
Why it improves the app: The app has 41 generated SVGs and not one box plot, despite
spec unit 2 requiring "תרשים קופסה: חציון, רבעונים, שפמים, חריגים". These three cover the
whole reading skill: comparison, skew, and the degenerate case that breaks naive reading.
All three are deterministic from listed data, so
scripts/figures.mjscan compute the quartiles rather than eyeball them. Use for: new figures (3) | quiz questions (~4)
p39–p41 (printed 27–29) — frequency, relative frequency, cumulative relative frequency
What is there: Table 1.10–1.12 build the three columns step by step from 20 students' work hours. Table 1.13 is 100 semi-professional soccer players' heights in 8 grouped intervals with all three columns filled. Examples 1.14–1.16 then ask "% less than 65.95" (23%), "% between 61.95 and 65.95" (18%), "% more than 65.95" (77%), "number between 61.95 and 71.95" (87). Why it improves the app: Spec unit 1 subtopic 5 (שכיחות, שכיחות יחסית, שכיחות מצטברת) — absent from the app. n = 100 makes every relative frequency a percentage with no division, which is ideal for a no-calculator drill. Also the natural place to introduce the "less than" vs "at most" vs "between" reading distinction. Use for: lesson prose | quiz questions (~5)
p61 (printed 49) — Exercise 39, twelve levels-of-measurement items
What is there: One exercise, twelve sub-items to classify as nominal / ordinal / interval / ratio: soccer ability (Superior/Average/Above average), baking temperatures, crayon colours, social security numbers, incomes in dollars, satisfaction coded 1/2/3, political outlook (extreme left … extreme right), time of day on an analog watch, distance in miles, the dates 1066/1492/1644/1947/1944, heights, letter grades A–F. Why it improves the app: A ready-made bank for a unit the app doesn't cover, and it contains three of the spec's named traps verbatim: a numeric code that is really nominal (SSN), a satisfaction rating that looks quantitative but is ordinal, and years/dates as interval (no absolute zero — you cannot say "twice as late"). "Time of day on an analog watch" is a fourth, harder case worth keeping as the hardest distractor. Use for: quiz questions (~12, one per item)
p129 (printed 117) — Chebyshev's Rule
What is there: Six lines. For ANY data set, whatever the shape: at least 75% within
2 SD, at least 89% within 3 SD, at least 95% within 4.5 SD. Immediately contrasted with the
Empirical Rule, which is flagged as holding only for bell-shaped symmetric data.
Why it improves the app: z-scores/lesson.md correctly warns that 68-95-99.7 holds only
under normality — and then leaves the student with nothing. Chebyshev is the answer to "so
what is still true?", and it converts a vague caveat into a usable bound. It directly arms
the spec trap "ציון Z גבוה לא אומר אחוזון גבוה אם ההתפלגות אינה נורמלית". Six lines of prose
for a real gap.
Use for: lesson prose (z-scores) | quiz questions (~2)
z-scores
p143–p146 (printed 131–134) — Practice items 64 and 68 (see full entry above)
Use for: correcting the "סימטרית → שלושתם מתלכדים" row of the app's skew table, and two "מה נובע" quiz items.
p146 (printed 134) — Exercise 71, Fredo vs Karl batting averages
What is there: Fredo bats .158 on a team averaging .166 (SD .012); Karl bats .177 on a team averaging .189 (SD .015). z = −0.67 vs z = −0.80. Why it improves the app: The app's flagship comparison example has both students above the mean. Comparing two negative z-scores ("who is less below average?") is a distinct sign-handling trap the app never poses, and the raw scores point the opposite way from the z-scores (Karl's raw average is higher). Two lines of data, a whole extra question type. Use for: quiz questions (~2)
p351 (printed 339) — Ex 6.4, two cohorts of Chilean male heights
What is there: 15–18-year-old Chilean males: 1984–85 cohort Y ~ N(172.36, 6.34); 2009–10 cohort X ~ N(170, 6.28). Then x = 160.58 and y = 162.85 both give z = −1.5. Why it improves the app: The app teaches "same raw score, different z" (Liron/Michael). This is the mirror image — different raw scores, same z — and it makes the "z measures position relative to its own distribution" point from the other direction. The spec names this trap: "שאלות המשלבות שתי התפלגויות עם ממוצעים וסטיות תקן שונים". Real published means. Use for: quiz questions (~2) | lesson example
confidence-intervals (the CLT / sampling-distribution section)
p388–p390 (printed 376–378) — Ex 7.9, individual vs mean, same cutoff
What is there: Cell-phone customers' excess minutes, a strongly right-skewed population
with μ = 22 and σ = 22. For a sample of n = 80: P(x̄ > 20) = 0.7919. For one randomly
chosen customer: P(x > 20) = 0.4029. Part (c) spells out why: "the probabilities are not
equal because we use different distributions for individuals and for means. When asked to
find the probability of an individual value, do not use the CLT."
Why it improves the app: This is the strongest single item in chapters 6–7. The app's
confidence-intervals/lesson.md explains SD vs SE beautifully but never shows the
consequence — that the same threshold has wildly different probabilities depending on
whether you are asking about a person or about an average. And because the population here is
skewed, it also nails the spec trap "CLT חל על ההתפלגות של הממוצע, לא על ההתפלגות של הנתונים":
for the individual you may not use a normal curve at all. Drop the exponential machinery,
keep μ = 22, σ = 22, n = 80, and the two probabilities.
Use for: worked problem | lesson prose | quiz questions (~3)
p405 (printed 393) — Exercise 67, three true/false statements
What is there: Verbatim: "When the sample size is large, (a) the mean of X̄ is approximately equal to the mean of X; (b) X̄ is approximately normally distributed; (c) the standard deviation of X̄ is approximately the same as the standard deviation of X." Why it improves the app: A ready-made three-option item whose false leg (c) is the SD/SE confusion the spec calls "ההבחנה הקריטית". Three statements, one wrong, no arithmetic — this is literally the מתא"ם multiple-choice format. Translate almost as-is. Use for: quiz question (1, high value) | flashcard
p406 (printed 394) — Exercise 76, predict before calculating
What is there: Attention span of two-year-olds, strongly skewed, mean 8 minutes, n = 60. Part (e): "Before doing any calculations, which do you think will be higher? (i) P(an individual attention span is under 10 minutes) or (ii) P(the average of the 60 children is under 10 minutes)? Explain why." Part (g): "Explain why the distribution for X̄ is not exponential." Why it improves the app: The prediction-first framing is exactly the reasoning skill the exam tests, and (g) forces the student to say out loud that averaging changes the shape, not just the spread. This is a better phrasing of the app's own CLT point than the app currently has. Use for: quiz questions (~2) | lesson prose
p404 (printed 392) + answer on p413 (printed 401) — Exercise 63, IRS Form 1040
What is there: Time to complete Form 1040: mean 10.53 hours, SD 2, distribution unknown, n = 36. (d) Would you be surprised if the 36 taxpayers averaged more than 12 hours? (e) Would you be surprised if one taxpayer took more than 12 hours? Answers supplied: (d) yes — probability almost 0; (e) no — probability 0.2312. Why it improves the app: Same individual-vs-mean contrast as Ex 7.9 but with the numbers and the verdicts already printed, and phrased as "would you be surprised" rather than "find the probability" — which is how a no-calculator exam has to ask it. Also a clean n = 36 ≥ 30 case with an unknown parent distribution. Use for: quiz questions (~2) | worked problem
p409 (printed 397) — Exercise 90, maternity stays
What is there: Mean hospital stay 2.4 days, SD 0.9, n = 80. (f) is it likely an individual stayed more than 5 days? (g) is it likely the average of the 80 was more than 5 days? (h) which is more likely? Why it improves the app: Part (h) makes the comparison the question rather than a follow-up. Third instance of the pattern with different numbers — enough to build a whole drill set of the type the spec asks for ("10 שאלות מה יקרה ל-SE"). Use for: quiz questions (~2)
p405 (printed 393) — Exercise 69, wedge-shaped income distribution
What is there: A country where mean salary is $2,000/yr with SD $8,000, n = 1,000. (d) "How is it possible for the standard deviation to be greater than the average?" (e) "Why is it more likely that the average of the 1,000 residents will be from $2,000 to $2,100 than from $2,100 to $2,200?" Why it improves the app: (d) is a conceptual item the app never touches — SD > mean is a signature of extreme right skew with a floor at zero, and students often assume it is impossible. (e) tests whether the student knows the sampling distribution is centred at μ and falls off symmetrically. Both are answerable with zero arithmetic. Use for: quiz questions (~2) | lesson prose
p391 (printed 379) — TRY IT 7.11, the Boeing 757 doorway
What is there: The 757 carries 200 passengers; doors are 72 inches. Men's heights: mean 69.0, SD 2.8. (a) What doorway height lets 95% of men enter without bending? (b) What doorway height gives a 0.95 probability that the mean height of 100 men is under it? (c) "For engineers designing the 757, which result is more relevant — part a or part b? Why?" Why it improves the app: The best conceptual question in chapter 7. The sampling distribution gives a numerically tighter, more impressive-looking answer — and it is the wrong tool, because doorways are used by individuals, not by group averages. This is the "מובהק ≠ רלוונטי" instinct applied to the CLT, and it is memorable enough to survive to the exam. Works fully as a discussion item without any calculation. Use for: worked problem | quiz question (1)
p391 (printed 379) — TRY IT 7.10(c) and Ex 7.11(b)
What is there: (i) Systolic BP of women 18–24: mean 114.8, SD 13.1, normal. "If the sample were four females and we did not know the original distribution, could the central limit theorem be used?" (ii) Broadway attendees, ages 14–61, mean 30.9, SD 9: "Is it likely that the mean age of a sample of 25 could be more than 50? ... However, it is still possible for an individual in this group to have an age greater than 50." Why it improves the app: (i) is the n ≥ 30 / normal-parent condition asked as a direct question — the app states the rule of thumb in a blockquote but never drills it. (ii) is the individual-vs-mean point in a single sentence, with an explicit age range (14–61) that makes "an individual over 50 exists; a mean of 25 people over 50 is impossible" concrete. Use for: quiz questions (~2)
p378 (printed 366) — the dice framing of the CLT
What is there: Roll 1 die, then 2, then 5, then 10, taking the mean each time.
"As the number of dice increases: (1) the mean of the sample means stays about the same,
(2) the spread gets smaller, (3) the graph appears steeper and thinner."
Why it improves the app: The app's clt-sampling-distribution.svg uses a skewed
reaction-time parent at n = 4 and n = 36. A dice version is a complementary figure: the
parent is perfectly flat (uniform 1–6), which makes "the shape of the parent does not
matter" far more visually shocking than a skewed parent does — and every value is exactly
computable, so scripts/figures.mjs can generate the exact distribution of the mean of
n dice with no simulation and no Math.random.
Use for: new figure idea
research-critique
p64–p65 (printed 52–53) — Exercise 76, the 1936 Literary Digest poll
What is there: A magazine mailed 10,000,000 postcards to people drawn from its own
subscriber list, automobile registrations, phone books and club memberships; ~2,300,000 were
returned; it predicted Landon would beat Roosevelt in a landslide. Sub-questions: (a) why that
frame was unrepresentative of 1936 America; (b) what the low response rate does to
reliability; (c) are these sampling or nonsampling errors; (d) Gallup's competing poll of
30,000 using quota sampling.
Why it improves the app: research-critique/lesson.md makes exactly this claim —
"מדגם מוטה של 10,000 גרוע ממדגם אקראי של 300" — and has no example to hang it on. This is
the canonical historical case, and it is better than the app's abstract version because the
biased sample is seven thousand times larger and still wrong. It also maps one-for-one onto
the app's own five-stage sampling-funnel.svg: the frame excluded the poor (no car, no
phone), then non-response filtered further. Add it as the named anchor for that whole section.
Use for: lesson prose | worked problem | quiz questions (~3)
p61 (printed 49) — Exercises 29–33, the stroke-patient software studies
What is there: A full four-question cluster on one scenario. Two studies of the same problem-solving software, each on 200 stroke patients, given as 2×3 tables: Study 1 — used program 142 improved / 43 no change / 15 worse; did not use 72 / 110 / 18. Study 2 — used 105 / 74 / 19; did not use 89 / 99 / 4. Then: 29. Which study is correct? 30. The first study was run by the company that made the software, the second by the American Medical Association — which is more reliable? 31. Both groups concluded the software works — is that accurate? 32. The company treats the two studies as proof the software causes mental improvement — is that fair? 33. Patients who used the software were also in an exercise program; those who did not use it were not. Does this change the validity of the conclusion? Why it improves the app: This is an אשכול already assembled: real-looking frequency data, a funding-bias question, a causal-overreach question, and an explicit confound revealed only in the last part. The design is self-selection ("used" vs "did not use", not assigned), so the answer chain is exactly the app's rule #1 ("מסקנה סיבתית + מערך שאינו ניסויי = זו הבעיה"). Exercise 33 is a textbook משתנה מתערב with a named alternative explanation, which is what the three-part answer template needs. The single best find in chapter 1. Use for: worked problem (full 3-part critique) | reading-results cluster | quiz questions (~4)
p65 (printed 53) — Exercise 77, education and crime across 47 states
What is there: Crime and demographic statistics for 47 US states, 1960, from the FBI
Uniform Crime Report. "One analysis of this data found a strong connection between education
and crime, indicating that higher levels of education in a community correspond to higher
crime rates. Which of the potential problems with samples could explain this connection?"
Why it improves the app: A real, published, counter-intuitive correlation whose resolution
is a third variable (urbanisation — cities have both more schooling and more reported crime),
plus a reporting artefact (better-educated communities report crime more). The app's
confounds problem bank is all constructed scenarios; this one is real and the "obviously
absurd" direction makes the third-variable move unavoidable. Also a clean instance of an
ecological correlation, which the app never mentions.
Use for: worked problem | quiz questions (~2) | confounds problem bank
p50 (printed 38) — TRY IT 1.22, the apple-juice taste test
What is there: Four flaws in one commissioned study of teens' favourite fruit juice: (a) the survey is commissioned by the seller of one of the brands; (b) only two brands are included; (c) participants can see the brand as the samples are poured; (d) the result is 25% prefer X, 33% prefer Y, 42% no preference — and Brand X advertises "Most teens like Brand X as much as or more than Brand Y." Why it improves the app: Four distinct problems — sponsor bias, artificially restricted alternatives, absence of blinding, and a technically-true-but-deceptive summary — in a scenario short enough to fit on a phone screen. (d) is the sharpest: 25 + 42 = 67% is "as much as or more", which is literally true and completely misleading, and Brand Y actually won. That statistical-spin flavour is missing from the app's critique bank, which covers overreach but not deliberate reframing. Use for: worked problem | quiz questions (~3) | problem bank entry ("ניסוח מטעה של תוצאה")
p46 (printed 34) — the vitamin E lurking-variable passage
What is there: Six sentences. "You recruit subjects and ask whether they regularly take vitamin E. Those who do exhibit better health on average. Does this prove vitamin E is effective? It does not. People who take vitamin E regularly often take other steps to improve their health: exercise, diet, other supplements, choosing not to smoke." Why it improves the app: The clearest short statement of the healthy-user confound I found, and it names four specific co-varying behaviours rather than saying "there might be a confound" — which is precisely the difference the app calls out between a strong and a weak critique answer ("הסבר חלופי צריך להיות ספציפי למחקר"). Useful as a model answer. Use for: lesson prose (model critique) | quiz question (1)
p47 (printed 35) — the McClung & Collins placebo finding
What is there: A quoted result from a real study of performance-enhancing drugs: "believing
one had taken the substance resulted in [performance] times almost as fast as those associated
with consuming the drug itself. In contrast, taking the drug without knowledge yielded no
significant performance increment." (McClung & Collins, J. Sport & Exercise Psychology
29(3):382–94, 2007.)
Why it improves the app: A finding where the belief beat the drug and the drug without
belief did nothing. That is a far stronger motivation for blinding and placebo controls than
any generic statement, and the app's confounds/research-critique material asserts the need
for סמיות without ever showing why it can dominate the treatment.
Use for: lesson prose | flashcard | quiz question (1)
p35–p36 (printed 23–24) — Ex 1.13, two biased samples with real numbers
What is there: ABC College, 10,000 part-time students, average spend on books. Sample 1 — convenience, 10 students from an organic chemistry class: $128, 87, 173, 116, 130, 204, 147, 189, 93, 153. Sample 2 — every fifth name from a list of senior citizens in P.E. classes: $50, 40, 36, 15, 50, 100, 40, 53, 22, 22. Sample 3 — one student randomly drawn from each of ten disciplines: $180, 50, 150, 85, 260, 75, 180, 200, 200, 150. The text argues sample 1 is biased high (science texts are expensive) and sample 2 biased low (interest courses), and that sample 3 is unbiased but small. Why it improves the app: Two biased samples that are biased in opposite directions, with computable means ($142.0, $42.8, $153.0). Because every number is there, this makes a proper worked problem: compute the three means, then explain why the sampling method, not the arithmetic, is what makes two of them wrong. The app's sampling-funnel section is entirely qualitative. Use for: worked problem | quiz questions (~2)
p32 (printed 20) — the "Critical Evaluation" checklist
What is there: Nine bulleted problems to check in any study: non-representative sample ·
self-selected sample · sample size · undue influence (asking questions in a way that
influences the answer) · non-response/refusal · causality from correlation · self-funded or
self-interest studies · misleading use of data (improperly displayed graphs, incomplete data,
lack of context) · confounding.
Why it improves the app: Two of the nine are not in the app's בנק הבעיות: loaded/leading
question wording, and sponsor/self-interest bias. Both are easy to check for and both appear
in real מתא"ם-style research descriptions. Two new rows for the problem-bank table.
Use for: two new rows in research-critique/lesson.md problem bank | quiz questions (~2)
p64–p65 (printed 52–53) — Exercises 74, 78, and 36, the "question wording / self-selection" trio
What is there: 74 — a "random survey" of 3,274 people born since 1971 reporting that 48% would spend $2,000 on computer equipment; the extra detail, revealed afterwards, is that the survey was filled out by visitors to a Smithsonian road show at the LA Convention Center and was reported by Intel. 78 — "Do you feel happy paying your taxes while some politicians are allowed to use loopholes and avoid paying their fair share of taxes?" — all 11 respondents answered NO. 36 — "Do you prefer the delicious taste of Brand X or the taste of Brand Y?" Why it improves the app: 74 is the "big n does not fix bias" point again but with a sponsor attached; 78 and 36 are one-line loaded questions that make perfect flashcard-length quiz items. Cheap to translate, three distinct flaws. Use for: quiz questions (~3)
confounds
p47–p48 (printed 35–36) — Ex 1.20, the scented-mask maze study
What is there: A real study (Smell & Taste Treatment and Research Foundation): subjects completed pencil-and-paper mazes three times wearing a floral-scented mask and three times wearing an unscented mask, with the order randomly assigned; time and the subject's impression of the scent were recorded. Questions: identify explanatory and response variables, name the treatments, identify lurking variables, and "is it possible to use blinding in this study?" The published answer is the good part: subjects obviously know whether they can smell flowers, so subjects cannot be blinded — but the researcher timing the maze can be. Why it improves the app: A within-subjects design with proper counterbalancing (spec units 16 and 25) plus the partial-blinding insight — that single-blind may be impossible on the subject side while still being achievable on the rater side. The app's problem bank lists "היעדר סמיות" as a single flat item; this shows it has two independent halves. Real study, clean design, six sentences. Use for: worked problem | quiz questions (~2) | refine the "סמיות" row of the problem bank
p48 (printed 36) — Ex 1.21, birth order
What is there: Four sentences. "A researcher wants to study the effects of birth order on personality. Explain why this study could not be conducted as a randomized experiment." Answer: you cannot randomly assign a person's birth order; without random assignment there will be differences between the groups other than the explanatory variable. Why it improves the app: The shortest, cleanest statement of the "משתנה נמדד ולא מתופעל ⟹ אין הסקה סיבתית" chain, phrased as a question rather than a rule. Birth order is a psychology variable, so it lands better than the app's generic examples, and it generalises into a whole item type: gender, smoking status, class, diagnosis, native language. Use for: quiz question (1) | seed for a "which of these can be randomly assigned?" set (~5)
p68 (printed 56) — Exercise 87, sleep deprivation and driving
What is there: "How does sleep deprivation affect your ability to drive? A recent study measured the effects on 19 professional drivers. Each driver participated in two experimental sessions: one after normal sleep and one after 27 hours of total sleep deprivation. The treatments were assigned in random order. In each session, performance was measured on a variety of tasks including a driving simulation." Why it improves the app: A five-line design description containing every marker the exam tests: within-subjects (same 19 drivers in both conditions), counterbalancing ("random order"), a within-subject df of 18, and a multiple-outcome hazard ("a variety of tasks" → השוואות מרובות). It is the ideal stimulus for the "identify the design and the right test" drill the spec wants 40 of (unit 38), because a single description supports four different questions. Use for: worked problem | choosing-a-test / t-tests quiz questions (~4)
reading-results
p69 (printed 57) — Exercise 89, the airline complaints bar chart
What is there: A bar chart of February 2013 US DOT passenger complaints for six airlines:
United ~125, American ~118, Delta ~48, Alaska/Pinnacle/Airtran ~5 each. The question:
"Alaska, Pinnacle, and Airtran have far fewer complaints reported than American, Delta, and
United. Can we conclude that American, Delta, and United are the worst airline carriers since
they have the most complaints?"
Why it improves the app: The count-versus-rate trap, and the app does not have it. Its
reading-results lesson covers percentage-base errors within a table but nothing about a raw
count presented where a rate is needed. The chart makes the wrong answer visually obvious,
which is what makes it a good trap: the correct answer requires noticing what is not on the
axis (passengers flown). One of the highest ratio of exam-relevance to effort in the book, and
the chart is trivial to regenerate with real numbers.
Use for: quiz questions (~2) | new figure | reading-results lesson section
p69 (printed 57) — Exercise 88, the Acme Investments ad
What is there: Two side-by-side line graphs captioned "As the graphs show, Acme
consistently outperforms the Other Guys!" The left panel's line climbs steeply, the right
panel's is nearly flat — because the two panels use different y-axis scales. The question asks
to describe the misleading visual effect and how to correct it.
Why it improves the app: A figure idea the app can execute better than the book:
generate both panels in scripts/figures.mjs from a single shared data series with two
different y-ranges, so the deception is provably about the axis and not the data. That is a
demonstration the textbook cannot make (its two series really are different) and it directly
serves spec unit 39. Fixed axis, no Math.random, deterministic.
Use for: new figure (high value) | quiz question (1) | reading-results lesson
p97–p98 (printed 85–86) — "How NOT to Lie with Statistics"
What is there: About a page of prose on graphical deception: pie charts with too many
slices or with two years set side by side (the total changes, so slice sizes are not
comparable); histograms with varying bar widths, where a wider category gets more area
and so reads as more probable; time-series axes whose time unit changes part-way (years, then
months); axes that do not begin at zero; and choosing units (pennies rather than thousands of
dollars) or long aggregation windows to smooth or exaggerate a trend.
Why it improves the app: reading-results teaches how to read a correct graph but never
how a graph lies. Three of these — unequal bar widths, a truncated axis, and a changed time
unit — are checkable in two seconds and would each make a one-line quiz item. Clearer and more
systematic than anything the app currently says about graphs.
Use for: lesson section (~½ page) | quiz questions (~4)
p197 (printed 185) — Ex 3.21, the 100 hikers table
What is there: A 2×3 table (Women/Men × Coastline / Lakes-and-streams / Mountain peaks)
with four cells and three totals left blank, to be completed from the margins. Then:
(b) are "being a woman" and "preferring the coastline" independent? — checked as
P(F AND C) = 18/100 = 0.18 versus P(F)·P(C) = 0.45 × 0.34 = 0.153, therefore not
independent. Then (c) a conditional, P(M | L) = 25/41, with the explicit note that "the
sample space for this problem is not all 100 hikers — it is the 41 who prefer lakes and
streams."
Why it improves the app: Three things at once, all missing: completing a partly-filled
two-way table from the margins (the same arithmetic as χ² expected frequencies,
E = row × column / N, spec unit 26); the formal independence check; and the "reduced sample
space" statement, which is the cleanest wording I found for the app's own denominator trap.
The app's reading-results percentage table is 2×2 with everything filled in — this is the
harder, more exam-like version.
Use for: worked problem | quiz questions (~3) | feeds the χ² unit
p66–p67 (printed 54–55) — Exercise 82, the broken frequency table
What is there: 19 immigrants' years lived in the US, presented as a frequency table with deliberate errors planted in it. (a) find the errors and explain how someone might have arrived at the wrong numbers; (b) "Explain what is wrong with this statement: '47 percent of the people surveyed have lived in the U.S. for 5 years.'"; (c) fix it; then five reading questions (five or seven years; at most 12; fewer than 12; five to 20 inclusive). Why it improves the app: (b) is the cumulative-versus-exact misreading — 0.4737 is the cumulative entry at 5, so the correct statement is "5 years or fewer". That is a real, recurring exam error and the app has nothing on it. Items (d)–(g) then drill the at-most / fewer-than / inclusive-range boundary distinctions, which is the other half of the same skill. A companion error-hunt exists at Ex 1.17, p43 (printed 31), where a frequency column sums to 18 instead of 19. Use for: worked problem | quiz questions (~4)
p196 (printed 184) — TRY IT 3.20, stretching and injury
What is there: 800 athletes. Stretch before exercising: 55 injured, 295 not (350 total). Do not stretch: 231 injured, 219 not (450 total). Column totals 286 / 514. Why it improves the app: The injury rate is 15.7% among stretchers versus 51.3% among non-stretchers — a huge, tempting difference in an observational study where athletes chose whether to stretch. It is a two-line table that supports a conditional-probability computation and a causal-overreach critique from the same numbers, which is exactly how an אשכול is built. The app's contingency-table example (workshop/pass) has a much weaker effect and no self-selection angle. Use for: worked problem | quiz questions (~3)
p69–p70 (printed 57–58) — Exercises 90 and 91
What is there: 90 — a distance-learning survey of 771 students reporting: 96% have a computer at home, 65% unable to come to campus, 24% aged 41+, 95% want more DL courses, 17% took DL due to a disability, 13% live 16+ miles away, 71% took DL for transfer requirements. Part (c): "If the same survey were done at Great Basin College in Elko, Nevada, would the percentages be the same? Why?" 91 — Martin, a college newspaper reporter, checks stock availability for one popular textbook in each of seven subjects at a random sample of major online retailers, and writes an article concluding about all college textbooks; the exercise asks for a written analysis of representativeness, sources of bias, and improvements. Why it improves the app: 90(c) is a bare תוקף חיצוני question with numbers attached (the percentages sum well past 100 because they are separate items — a small extra trap). 91 is structurally a מתא"ם open critique question: a stated conclusion that is far broader than the sample supports, with an obvious sampling flaw (one textbook per subject, only large retailers). Both are ready to translate with minimal editing. Use for: research-critique problem bank (2 scenarios) | quiz questions (~2)
Optional aside, not recommended for development
p48–p49 (printed 36–37) — the Diederik Stapel case. A Dutch social psychologist whose fraud tainted 55 papers and 10 supervised PhDs; the investigating committee found he created datasets that confirmed prior expectations, altered existing data, changed measuring instruments without reporting the change, and misrepresented the number of subjects. The report notes "statistical flaws frequently revealed a lack of familiarity with elementary statistics" and that co-authors who did not know statistics well simply trusted him. The adjacent paragraph mentions researchers who stop collecting data once they have just enough to prove what they hoped — which is optional stopping, a real item in the "ניתוח והיסק" category of the app's problem bank. It is a psychology case and it is memorable, but there is no ethics unit in the 44-unit spec, so develop it only if you decide to add one. The optional stopping sentence alone could be folded into the existing problem bank as a one-line entry next to "השוואות מרובות".
OpenStax Introductory Statistics 2e — chapters 8, 9, 10
Mined against the existing app content (confidence-intervals, hypothesis-testing,
power-and-errors, t-tests, choosing-a-test) and content/_docs/content-spec.md.
Page convention: PDF page number first, printed book page in parentheses.
Headline judgement. Chapter 8 is almost entirely redundant — the app's
confidence-intervals lesson is more precise than the textbook and its worked problems
already cover the same ground with cleaner numbers. Chapter 9's conceptual exercise
sets are the best material in scope. Chapter 10's body text is actively wrong for this
exam (see SKIPPED), but its scenario banks are the single most useful thing in all
three chapters.
The one real content gap found: the one-sample -test has no lesson section and no
worked problem anywhere in the app — it exists only as a row in choosing-a-test
tables and as the paired-test reduction. Spec unit 14 (⭐B) asks for it explicitly.
OpenStax 9.5 fills it.
SKIPPED — examined and rejected
Proportions: all of it. 8.3 (p431–p438 / 419–426) and 10.3 (p535–p538 / 523–526), plus Examples 9.17, 9.18, 9.19 in 9.5. Checked the content spec first, as instructed. There is no unit on proportion inference among the 44. The test catalogue in unit 38 (line 1091) is: one-sample · independent · paired · Pearson · Spearman · linear regression · one-way ANOVA · two-way ANOVA · within/mixed ANOVA · chi-square. No -test for a proportion, no two-proportion test. Categorical data on this exam goes through (unit 26), which is chapter 11, not chapter 8/10. The "plus four" interval (p434–p436 / 422–424) is a US-textbook idiosyncrasy that appears in no Israeli syllabus. Do not build content from these pages.
10.1 body text and formula apparatus, p524–p525 (512–513) and Formula Review p548–p549
(536–537). This is a trap for anyone mining chapter 10 naively. OpenStax teaches the
unpooled Aspin-Welch test as the default, with the Satterthwaite df formula, and says
in bold "Do not pool the variances." The app teaches pooled Student with
— which is what the exam requires, and what the app's own lesson,
flashcards and tt-p01 are built on. Mining these pages for formulas would contradict the
app and introduce a df formula the exam never uses. Take the scenarios from chapter 10,
never the formulas. (The one exception is , which OpenStax gives only inside
Cohen's — see the power-and-errors entry below.)
10.2 Two Population Means with Known σ, p532–p534 (520–522). The two-sample -test. Not in the spec's catalogue; knowing both population σ's is a textbook fiction. Examples 10.6 (floor wax) and 10.7 (senators) rejected.
8.1 body, p419–p431 (407–419). Interval structure, effect of confidence level,
effect of , working backwards from an interval, and the sample-size formula
are all already in confidence-intervals/lesson.md and
ci-p01–ci-p03, with rounder numbers and a sharper split between centre and width.
Examples 8.1–8.7 add nothing. (One narrow exception survives — see the CI section.)
8.2 body, p430–p431 (418–419). The Gosset/Guinness history and the properties of the distribution are covered by the app's "מתי ומתי " section, which is tighter.
Chapter Reviews, Key Terms, Formula Reviews — p444–p446 (432–434), p498–p500 (486–488), p548–p549 (536–537). Pure restatement; the app's flashcard decks already carry these definitions in Hebrew.
All Stats Labs and Collaborative Exercises — 8.4/8.5/8.6 (p439–p443 / 427–431), 9.6 (p495 / 483), 10.5 (p544–p547 / 532–535), Collaborative Exercise p476 (464). Classroom activities involving newspapers and M&Ms.
All TI-83/84 boxes, per brief.
Example 9.10, p481 (469) — "If the -value is low, the null must go." An English rhyme; does not survive translation.
Historical Note, p485 (473). Reconciles the critical-value route with the -value
route on one graph. The app already says this twice ("שתי הדרכים נותנות תמיד את אותה
הכרעה") and p-value-vs-alpha.svg already shows the two areas. Only the figure idea
survives — listed below.
Table 9.3 distribution-selection table, p478 (466). The app's choosing-a-test
decision tree is richer.
Homework #83–86, p555 (543) — depend on Appendix C data files not in scope.
t-tests
p487 (475) + p493 (481) — Examples 9.16 and 9.20: one-sample from raw data
What is there: Two complete one-sample -tests worked from small raw datasets.
Ex. 9.16: ten first-test statistics scores (65, 65, 70, 67, 66, 63, 63, 68, 72, 71),
, , , , one-tailed , reject at
. Ex. 9.20: eleven glass-conductivity measurements (NIST), ,
, , , — presented in an explicit four-step frame
(State the Question / Plan / Compute / State the Conclusion). Both explicitly reason
"no population standard deviation is given, therefore Student's ".
Why it improves the app: This is the app's clearest hole. t-tests/lesson.md is
titled "מבחני לשני מדגמים" and covers only independent and paired; there is no
one-sample worked problem in any topic. Spec unit 14 (⭐B) demands "מבחן למדגם
יחיד מקצה לקצה" plus 8 questions. Both datasets are small enough to compute by hand in
Hebrew step-by-step, and both means/SDs verify exactly.
Use for: worked problem (2) · lesson prose (a new "מדגם יחיד" section in
t-tests/lesson.md) · quiz questions (~4, including "σ not given ⟹ , ")
p540–p541 (528–529) — Example 10.11: paired , hypnotism and pain
What is there: Eight subjects, sensory pain rating before and after hypnotism.
Difference scores {0.2, −4.1, −1.6, −1.8, −3.2, −2, −2.9, −9.6}, ,
, , , one-tailed . The write-up spends a full
paragraph justifying the direction of vs from what
"improvement" means on this scale, and notes in bold that the test is now a test of a
single population mean.
Why it improves the app: The app's paired problems (tt-p02, tt-p03) both use
tidy, uniformly-signed differences where every subject moves the same way. Here one
subject moves the wrong way (+0.2) and one moves enormously (−9.6), which is what real
before/after data looks like — and it is a psychology-domain study, unlike the app's
generic anxiety questionnaire. The explicit reasoning about which direction counts as
"improvement" on a reverse-scored measure is a distinct trap the app never poses.
Use for: worked problem (1) · quiz questions (~2 on direction of when a low
score is the good outcome)
p542–p543 (530–531) — Example 10.12: paired wrecked by one outlier ()
What is there: Four softball players' max lift before and after a strength class.
Differences {90, 11, −8, −8}: , , ,
, — not significant. OpenStax then adds an unusually candid NOTE:
the data are right-skewed, 90 is probably an outlier, it alone drags the mean positive
while three of the four differences are negative, is far too small, and the
normality assumption on the differences should be confirmed.
Why it improves the app: This is a single scenario that fires four of the app's
recurring traps at once — mean ≠ typical value, tiny ⟹ low power ⟹ uninformative
non-significance, assumption violation, and "לא נמצא הבדל מובהק ≠ הוכח שאין הבדל". The
app teaches each of these separately and abstractly; nothing in it shows all four
colliding on four real numbers. Excellent research-critique material (spec unit 41).
Use for: worked problem (1, cross-linked to power-and-errors and
research-critique) · quiz questions (~3) · new figure idea (below)
p543–p544 (531–532) — Example 10.13: shot-put, dominant vs weaker hand,
What is there: Seven eighth-graders push a shot-put with each hand. Differences
{2, 12, 7, −1, 2, 0, 4}, , , , ,
two-tailed → do not reject at .
Why it improves the app: A genuinely borderline result where and the mean
difference looks substantial, yet the two-tailed test fails — and the one-tailed would
be , i.e. significant. This is exactly the structure of the app's ht-p02
(same data, three tailedness scenarios) but here with real paired data, so it
simultaneously drills the paired-design identification and the one- vs two-tailed trap.
The app's ht-p02 uses invented numbers; this is better.
Use for: worked problem (1) · quiz questions (~2)
p553 (541) — Practice 10.4 #63–77: three tiny paired datasets
What is there: Three before/after tables with single- and double-digit integers. (a) Software patch, 8 installations, crashes before/after. (b) Juggling class, 6 subjects, balls juggled before/after — differences {1,1,3,2,1,2}, , , exactly, , significant at . (c) Blood-pressure medication, 6 patients — differences {−3,−3,+1,−2,+1,+... } give , , , , not significant at even though four of six patients improved. Why it improves the app: Arithmetic clean enough to do in the head — the juggling one lands on exactly. The blood-pressure/juggling pair is a ready-made contrast: same , same design, same , opposite decisions. All three are good "how many difference scores? what is ?" items, which is the app's stated most-common paired error. Use for: quiz questions (~5) · one alternative worked problem
p557–p558 (545–546) — Homework #98 + Table 10.29: marital-satisfaction paired data
What is there: Ten dual-career couples rate "I'm pleased with the way we divide responsibilities for childcare" on a 1–5 scale, husband and wife separately. Testing whether the mean husband−wife difference is negative. Differences: {0,0,−2,0,−2,−1,0,0,0,0}, , , , , one-tailed . Why it improves the app: A genuine psychology paired design where the "pairs" are two different people matched by relationship, not one person measured twice — the app's lesson lists "זוגות מותאמים" in its identification table but has no worked example of it. Six of ten differences are exactly zero, which makes vivid why (not , ) is the denominator. Lands one-tailed just under 0.05, so it also exercises the significance-boundary discussion. Ordinal 1–5 ratings also open a legitimate "is a -test appropriate on a 5-point scale?" critique question (spec unit 1 trap). Use for: worked problem (1) · quiz questions (~2)
choosing-a-test
p549–p550 (537–538) — Practice 10.1 #1–15: the six-way test-selection bank
What is there: Fifteen one- or two-sentence research scenarios, each to be classified
into one of six labelled options: (a) independent means, σ known · (b) independent means,
σ unknown · (c) matched or paired · (d) single mean · (e) two proportions · (f) single
proportion. Several are built traps: #3 — "ten windshields are tested without the
treatment. The same windshields are then treated and the experiment run again"
(sounds like two groups; is paired). #14 — a researcher tests 12 routers' native range,
then re-tests the same routers with a booster (paired). #10 — eight subjects'
sleep recorded before and after a medicine (paired). #4 and #8 slip in a known
population σ. #5, #6, #12 describe only one group, so they are single-sample.
Why it improves the app: This is the highest-value find in chapter 10. choosing-a-test
has 25 quiz items and t-tests has 23; the spec asks for "10 שאלות זיהוי מערך מתוך תיאור"
(unit 16) plus "6 שאלות בחירת מבחן" (unit 38). The app's lesson explicitly names the
paired-disguised-as-independent trap but is thin on items that actually spring it. These
are already in the exact multiple-choice shape the app's quiz.md uses — one option set,
fifteen stems — so they translate almost mechanically. Drop the two proportion options and
substitute the app's own catalogue (Pearson / χ² / ANOVA) as distractors.
Use for: quiz questions (~10 in choosing-a-test, ~3 in t-tests)
p564–p566 (552–554) — Bringing It Together #124–135: twelve more selection scenarios
What is there: The same six-option format, twelve further scenarios, plus two items
in explicit 4-option MCQ form: #134 — 41 left-handed vs 41 right-handed preschoolers
on motor competence, means 97.5 and 98.1, SDs 17.5 and 19.2, "determine the appropriate
test and best distribution"; #135 — four golfers' 18-hole scores before and after a
technique class, four options including "a test of two independent means".
Why it improves the app: A second, independent batch, so there is enough here to cover
both choosing-a-test and t-tests without reusing stems. #134 is doubly useful: it is a
real psychology study and the numbers make it obvious the difference is nowhere near
significant with SDs of ~18 — so the same stem serves an effect-size / power question.
#135 with pairs pairs naturally with Example 10.12 above.
Use for: quiz questions (~6) · one power-and-errors stem
power-and-errors
p501 (489) — Practice 9.2 #11–20: Type I / Type II errors stated in context
What is there: Ten short scenarios, each requiring the two errors to be stated as full
sentences in the terms of the scenario, plus three that go further:
#16 "which is the error with the greater consequence?" (surgeons deciding whether to
operate); #20 the same question after the null is flipped ("the sample contains
E-coli" vs "does not contain E-coli") — showing that which error is worse depends entirely
on how was phrased; #17 power = 0.981, find ; #19 given
and , find the power.
Why it improves the app: The app's 25 power-and-errors items are overwhelmingly
structural — definitions, the 2×2 matrix, what happens to power when grows. Only
pw-p03 asks "which error is more serious", once. This supplies a whole battery of
translate-the-scenario items, and the #19/#20 pair adds something the app genuinely lacks:
that can legitimately be phrased in the negative direction, which swaps which error
is which. That is a real exam-style reversal.
Use for: quiz questions (~6) · lesson prose (a short "מי משתי הטעויות חמורה יותר תלוי
בניסוח " note)
p477–p478 (465–466) — Examples 9.6–9.8 and Try It 9.7–9.8: the same, in richer scenarios
What is there: Four fully worked in-context error pairs — an emergency crew judging
whether an accident victim is alive (Type I is worse: they won't treat); Genetic Labs
claiming to raise the chance of a male birth (Type I is worse: people will buy it); an
experimental cancer drug claiming a ≥75% cure rate (Type II is worse: patients will choose
it); a marine-fisheries toxin threshold for banning clam harvesting. Try It 9.8 is
already formatted as a four-option multiple choice — "identify the Type I and Type II
errors from these four statements" — with the four statements being the four cells of the
2×2 matrix in prose.
Why it improves the app: Try It 9.8's format is exactly quiz.md's. The four scenarios
differ in which error is worse and why, which forces the reasoning rather than a
memorised answer. The cancer-drug one (Type II worse) is a useful counterweight — most
people default to "Type I is always the serious one" because of the criminal-trial analogy
the app opens with.
Use for: quiz questions (~5) · lesson prose (a small table of four scenarios × which
error is worse, to sit under the existing 2×2)
p505 (493) — Homework #68–71: Type I/II identification, already in 4-option MCQ form
What is there: Four items where the four options are near-identical sentences differing only in which state of the world is real. #68 is the best: = "the drug is unsafe", what is the Type II error? Options: (a) conclude safe when unsafe (b) not conclude safe when safe (c) conclude safe when safe (d) not conclude unsafe when unsafe. #70 and #71 do the same for "at least seven hours of sleep" and "mean phone hours is higher". Why it improves the app: Ready-made distractor sets. #68's null is phrased negatively — so the intuitive answer ("Type II = missing a real danger") is wrong, and you must actually apply the definition. That is precisely the app's stated exam angle. Also note #70/#71's distractors turn on "at least / at most / more than", which stresses the equality-side rule. Use for: quiz questions (~4, near-verbatim translations)
p530–p531 (518–519) — Cohen's table + Examples 10.4 and 10.5: computed from real tests
What is there: with spelled out
(this is the only place OpenStax uses pooling), the 0.2/0.5/0.8 table, and then
computed for two tests already worked earlier in the chapter. Ex. 10.4: two colleges'
math-course counts, and , means 4 and 3.5, SDs 1.5 and 1 →
(not significant) but (small-to-medium). Ex. 10.5: online vs face-to-face
final exams, each → (significant) and (large).
Why it improves the app: power-and-errors/lesson.md ends on a 2×2 table of
significance × effect size with no numbers in it at all, and the app has no worked
Cohen's computation anywhere. These two examples populate two of the four cells with
real, verifiable arithmetic ( checks out). Pairing them lets a
Hebrew worked problem compute and from the same data and show they answer
different questions — which is the whole point of the unit.
Use for: worked problem (1) · lesson prose (fill the 2×2 with these two cases) ·
quiz questions (~2)
hypothesis-testing
p504–p505 (492–493) — Homework #62–65: stating /, with parameter-vs-statistic traps
What is there: #62 gives ten one-line claims to convert into hypothesis pairs, ranging over "the mean is 34", "at most 60%", "at least =\le\ge\ne<>. **#65** is the valuable one: a 4-option MCQ where the distractors are (a) H_0: \bar{x} = 4.5H_0: \mu = 4.75\ge<H_0: \mu = 4.5,\ H_a: \mu > 4.5H_0 always contains equality, but never states that **hypotheses are about parameters, never about statistics** — and #65's distractor (a) is exactly that error. Combined with Table 9.1 (p474 / 462), which lays out the three legitimate symbol pairings (=\ne\ge<\le>=\ne\mu\bar{X}$; plus the three-pairing table)
p556 (544) — Homework #87–92: conclusion-wording multiple choice
What is there: #88 and #89 (and #92) present four near-identical conclusion sentences
that differ only in "sufficient evidence" vs "insufficient evidence" and in the direction
claimed, and ask which one correctly states the outcome. #89's four options cross
sufficient/insufficient with "night students better" / "day students better" / "there is a
difference".
Why it improves the app: The app's ten invalid phrasings target the meaning of
; these target the conclusion sentence, which is a different failure mode and the
one the exam's ביקורת questions actually present. The distractor pattern — swap
"insufficient evidence that there IS a difference" for "sufficient evidence that there is
NO difference" — is the app's most-emphasised trap, here in ready-made option form.
Use for: quiz questions (~4) · feeds reading-results too
p479 (467) — 9.4 "Rare Events": the birthday-party bubble
What is there: 200 plastic bubbles in a basket, exactly one contains a 100 bill. . Ali concludes the two of them were told wrong — there must be more than one p$ — ההגדרה המדויקת")
p506 (494) — Homework 9.5 #74–80: one-sample word problems, several with a σ/s trap
What is there: Seven one-sample test scenarios. #74 gives both a known
and a sample — and the earlier Practice #46 (p503 / 491)
asks outright "since both σ and are given, which should be used, and why?".
#78 gives eight raw sick-day counts (12, 4, 15, 3, 11, 8, 6, 8) for a two-tailed test
against . #75, #76, #77, #79 vary from 12 to 81 with known or not.
Why it improves the app: The "both given — which do you use?" item is a clean,
compact trap that the app's confidence-intervals lesson prepares for conceptually
(Z vs t) but never poses as a test-selection question. #78's eight integers are small
enough for a full Hebrew hand computation and give the app a second one-sample dataset.
Use for: quiz questions (~3) · one worked problem variant
p485 (473) Figure 9.7 — new figure idea: all four objects on one axis
What is there: A single normal curve carrying, simultaneously, the shaded
region, the shaded -value region, the critical value () and the test statistic
(), with the visual point that the test statistic being further out than the
critical value is the same fact as .
Why it improves the app: The app has ht-rejection-regions.svg (regions only) and
p-value-vs-alpha.svg (two areas, two panels). Neither puts the critical value and
the test statistic and both areas on one axis, which is the picture that makes
"שתי הדרכים נותנות תמיד את אותה הכרעה" self-evident rather than asserted. Cheap to
generate in scripts/figures.mjs — one curve, two shaded tails at different cutoffs, two
vertical markers.
Use for: new figure idea
confidence-intervals
p429–p430 (417–418) and p431–p432 (419–420) — Examples 8.8 and 8.9: -intervals from raw data
What is there: Two complete -based confidence intervals computed from raw samples.
Ex. 8.8: fifteen acupuncture sensory ratings, , , ,
, , CI . Ex. 8.9: Human Toxome Project —
number of the 430 targeted industrial chemicals detected in each of twenty newborns' cord
blood, , , , , ,
90% CI .
Why it improves the app: confidence-intervals/problems.md has five worked problems
and all five use . The lesson explains that is the common case in practice and
gives the width comparison, but no problem ever builds a interval. Ex. 8.9 also uses a
90% level, so rather than the reflexive 2-point-something — good
against the habit of reaching for 1.96. Both datasets verify (8.9: sum 2549 / 20 = 127.45).
Use for: worked problem (1, using Ex. 8.9 — twenty integers, 90% level)
p447 (435) — Practice 8.1 #10–12: the three-way trade-off with margin of error held fixed
What is there: A Census survey, , . Then: #10 "if the Census wants to increase its level of confidence and keep the error bound the same, what changes should it make?" #11 "if it kept the error bound the same and surveyed only 50 instead of 200, what would happen to the level of confidence?" #12 "if it needed 98% confidence, would it have to survey more people?" Why it improves the app: The app's centre/width table treats , , and CL as four independent knobs and asks what each does to the width. These questions fix the width and ask what must give — which is strictly harder and is the form the question takes when it appears as a research-design question rather than an arithmetic one. Three items, no new theory needed, directly reusable against the app's existing identity. Use for: quiz questions (~3) · lesson prose (one line: "אם רוצים להעלות ביטחון בלי להרחיב את הרווח — חייבים להגדיל את ")
p446 (434) and p448 (436) — Practice #1 and #13: σ and both supplied
What is there: The elephant-calf item states (known) and from the sample; the lettuce item states (known) and . In both, the correct move is to use and the normal distribution and ignore entirely. Why it improves the app: A one-line trap. The app's Z-vs- table is keyed on "σ ידועה / אינה ידועה", which handles the case where only one of the two is given; it never poses the case where both appear on the page and the student must pick. Cheap to add, and it is the kind of distractor an exam writer reaches for. Use for: quiz questions (~2)
New figure ideas (for scripts/figures.mjs)
-
All four objects on one axis — from Figure 9.7, p485 (473). One normal curve with shaded, the -value shaded, the critical value marked and the test statistic marked, showing that "מעבר לערך הקריטי" and "" are one fact. Fills a real gap between the app's two existing hypothesis-testing figures. (Detailed above.)
-
The outlier that owns the mean — from Example 10.12, p542–p543 (530–531). A dot plot of four difference scores {90, 11, −8, −8} with marked. Three of four dots sit left of zero; the mean sits far right of every dot but one. Makes "ממוצע ההפרשים אינו ההפרש האופייני" visible in one glance, and supports the small- / assumption-violation discussion the app currently carries only in prose. Deterministic, trivial to generate, and no equivalent exists among the app's 41 figures.
Summary count
18 accepted items: 6 for t-tests, 2 for choosing-a-test, 4 for power-and-errors,
4 for hypothesis-testing, 3 for confidence-intervals (one, the figure note, doubles).
Roughly 55–65 translatable quiz questions, 8 candidate worked problems, 5 lesson-prose
additions, 2 figure ideas.
Weight of value by chapter: chapter 9 exercise sets ≫ chapter 10 scenario banks > chapter 10 paired examples > chapter 9 worked -tests ≫ chapter 8 (nearly nothing).
OpenStax Introductory Statistics 2e — chapters 11, 12, 13
Mined against the app's existing correlation, regression, anova-oneway,
choosing-a-test and reading-results content and against content-spec.md
units 17, 19, 20, 21, 22, 23, 26, 38, 39.
All page numbers are PDF pages, printed book page in parentheses (PDF = printed + 12).
SKIPPED — examined and rejected
11.4 Test for Homogeneity (p590–p594, printed 578–582) — out of scope.
Unit 26 in content-spec.md names exactly two chi-square tests: מבחן התאמה
(goodness-of-fit) and מבחן אי-תלות (independence), with df = k−1 and
df = (r−1)(c−1). Homogeneity is not on the מתא"ם and OpenStax teaches it with a
different df rule (df = columns − 1) that would actively contradict what the
app already teaches. Do not import it. The one salvageable item is Table 11.19
(men/women × living arrangements, 2×4) which can be re-framed as an ordinary
independence table — but the tables recommended below are better, so skip it.
11.6 Test of a Single Variance (p594–p596, printed 582–584) — out of scope. No unit in the spec mentions a test on a variance. Nothing here.
13.4 Test of Two Variances (p702–p704, printed 690–692) — out of scope. Not in the spec. OpenStax itself writes: "Many texts suggest that students not use this test at all, but in the interest of completeness we include it here." That is a clear signal. All of exercises 75–79 (p719–p721) go with it.
Example 11.4 — two coins (p585, printed 573). Rejected. The pedagogy is
confusing rather than instructive: OpenStax silently redefines the random
variable from four outcomes to three (number of heads), so df = 2 instead of 3.
It is arguably the wrong analysis and would teach the wrong reflex.
Ex 85 — obesity percentages (p612, printed 600). Rejected. Runs a chi-square on percentages rather than counts, which is a genuine methodological error. Not worth importing even as a trap; the app has cleaner critique material.
12.1 Linear Equations (p629–p632, printed 617–620), and Ex 1–14, 57–58, 62–66 (p663–p665, p670–p672). Rejected — pre-algebra slope/intercept drills (SCUBA rental fees, credit-card late fees, "is y = 10 + 5x − 3x² linear?"). Far below the app's level.
12.2 Scatter Plots (p632–p635, printed 620–623). Rejected. Four tiny
datasets (vocabulary by age, jump-shot practice, faculty/students, tuition vs
salary) and three unlabelled figures. The app's scatter-r-grid.svg,
r-zero-curved.svg and the new Anscombe section already do this job better than
anything on these pages.
CPI over time (p652–p653), life expectancy (p672), Olympic swim times (p674), Cuba PPP (p671), flu cases (p664–p665), Detroit homicide (p678). Rejected as a class — all time-series regressions. Year-as-X invites exactly the causal reading the app spends a whole section warning against, and none of them is small enough to compute by hand.
All TI-83/84 boxes, all Stats Labs (11.7, 11.8, 12.7–12.9, 13.5), all Collaborative Exercises, and all Chapter Review / Formula Review sections. Rejected per brief. The Chapter Review sections in particular (p601–p604, p662–p663, p707–p708) restate content the app already states better in Hebrew.
Ex 51–56 outlier concept questions (p670, printed 658). Rejected as redundant — the same "r went from 0.69 to 0.98 after removing a point" idea arrives with real numbers in Example 12.12 below.
chi-square — feeds choosing-a-test (and would justify its own topic)
This is where chapters 11 pays off most. The app currently has one chi-square worked example, a goodness-of-fit with three equal expected counts of 40. Unit 26 asks for "10 טבלאות לחישוב E ו-df". Everything below fills that.
p590 (printed 578) — 11.3 Example 11.7, anxiety level × need to succeed
What is there: A 3×5 contingency table, N = 400, with row and column totals printed, and two worked expected-count computations ( and ). Verified visually — the table renders cleanly with all marginals.
| Need to succeed \ Anxiety | High | Med-high | Medium | Med-low | Low | Row total |
|---|---|---|---|---|---|---|
| High need | 35 | 42 | 53 | 15 | 10 | 155 |
| Medium need | 18 | 48 | 63 | 33 | 31 | 193 |
| Low need | 4 | 5 | 11 | 15 | 17 | 52 |
| Column total | 57 | 95 | 127 | 63 | 58 | 400 |
Why it improves the app: This is the single best table in the three chapters
for this app. (a) It is a psychology scenario — anxiety and achievement
motivation — so it needs no re-skinning for a Hebrew psych-admissions exam.
(b) N = 400 is round, so every expected count is a clean division. (c) The
expected counts are wildly non-uniform (22.09 down to 8.19), which is exactly
what the app's equal-counts example fails to teach. (d) df = (3−1)(5−1) = 8 is
a perfect trap: the two seductive wrong answers, 15 (N-flavoured, i.e.
cells − 1) and 4, are both plausible. Unit 26 calls df "הטעות הכי נפוצה" and says
to drill it to automaticity — this is the table to drill it on.
Use for: worked problem (full: E for a named cell, df, χ², decision) ·
quiz questions (~5: df, two expected counts, which cell contributes most,
what a significant result licenses) · lesson prose replacing the current
equal-counts-only treatment
p616 (printed 604) — Ex 99, age group × net worth of young entrepreneurs
What is there: A 2×3 table, N = 40, both row totals equal to 20.
| Age \ Net worth (M$) | 1–5 | 6–24 | ≥25 | Row total |
|---|---|---|---|---|
| 17–25 | 8 | 7 | 5 | 20 |
| 26–30 | 6 | 5 | 9 | 20 |
| Column total | 14 | 12 | 14 | 40 |
Why it improves the app: Every expected count is an integer — 7, 6, 7, 7, 6, 7 — because both row totals are 20 and N = 40. I computed the whole test: , , not significant. That means the app can carry a complete, exactly-checkable chi-square worked problem where the user never touches a decimal, which is the only kind that survives being done in the head on a phone. It also delivers a non-significant chi-square, which the app currently lacks entirely (its only example rejects), so it can teach "χ² near its df means observed ≈ expected". Use for: worked problem (the flagship hand-computable one) · quiz questions (~3)
p584 (printed 572) — 11.2 Try It 11.3, number of pets, goodness-of-fit
What is there: A GOF against a given non-uniform distribution. Expected percentages 18 / 25 / 30 / 18 / 9; observed frequencies from n = 1,000 students: 210, 240, 320, 140, 90.
Why it improves the app: Expected counts are 180, 250, 300, 180, 90 — all
integers — and every contribution is exact:
,
. The app's goodness-of-fit story stops at "no preference ⇒ equal
counts". This is the other half: the null can be any stated distribution, and
then the expected counts differ per cell. Unit 26 sub-topic 2 ("שכיחויות נצפות
מול צפויות") is only half-taught without it. Bonus: one cell contributes 0 and
another contributes 8.89, which makes the "תרומת כל תא" sub-topic visible.
Use for: worked problem · new figure idea (bar chart of observed vs expected
with a non-flat expected line — the app's current
chi-square-observed-expected.svg shows a dashed horizontal expected line,
which quietly implies expected counts are always equal; a second figure with a
stepped expected line would correct that)
p582–p583 (printed 570–571) — 11.2 Example 11.3, streaming services
What is there: The same shape as above but fully worked. Expected percentages 10 / 16 / 55 / 11 / 8 of n = 600 → expected 60, 96, 330, 66, 48; observed 66, 119, 340, 60, 15. , . Crucially, an explicit boxed note: "df ≠ 600 − 1".
Why it improves the app: Two things. First, the cell contributions are
0.60, 5.51, 0.30, 0.55, 22.69 — one cell supplies 77% of the statistic. That
is a far sharper demonstration of "התרומה הגדולה מגיעה מהתא הכי חורג" than the
app's current 5.63 / 0.10 / 4.22 split, and it sets up the natural follow-up
question ("which category is the far west unlike the country on?"). Second, the
df ≠ n − 1 note names the exact confusion the spec flags — df comes from the
number of categories, never from the sample size. That deserves a sentence of
its own in the lesson.
Use for: lesson prose (the df ≠ n − 1 warning) · quiz questions (~3) ·
worked problem
p587 (printed 575) — 11.3 Example 11.5, why E = (row × column) / N
What is there: Before any table, OpenStax derives the expected-count formula from independence itself: if A and B are independent then P(A ∩ B) = P(A)·P(B), so with 755 drivers of whom 70 got a speeding ticket and 305 used a phone, . Try It 11.5 repeats it with clean numbers: 300 students, 50 music students, 97 on the honor roll → .
Why it improves the app: This is an explanation clearer than the app's
current phrasing. choosing-a-test/lesson.md gives the formula
(סכום שורה × סכום עמודה) / N as a rule to memorise. The derivation shows it is
not a rule at all — it is the definition of independence from unit 6, re-used.
That connects unit 26 back to unit 6 for free and makes the formula
un-forgettable rather than memorised. Two or three sentences of lesson prose.
Use for: lesson prose · flashcard
p588–p589 (printed 576–577) — 11.3 Example 11.6, volunteer type × hours
What is there: A 3×3 table with both the observed and the fully computed expected table side by side — the only place in the chapter where a complete expected table is printed.
| Volunteer type | 1–3 h | 4–6 h | 7–9 h | Row total |
|---|---|---|---|---|
| Community college students | 111 | 96 | 48 | 255 |
| Four-year college students | 96 | 133 | 61 | 290 |
| Nonstudents | 91 | 150 | 53 | 294 |
| Column total | 298 | 379 | 162 | 839 |
Expected: 90.57 / 115.19 / 49.24 · 103.00 / 131.00 / 56.00 · 104.42 / 132.81 / 56.77. , , p = 0.0113.
Why it improves the app: The value is the side-by-side pair of tables, which is a figure the app does not have and should: a 3×3 grid where each cell shows O above E, shaded by . That makes "where the relationship breaks" a picture rather than a sentence. N = 839 is ugly for hand arithmetic, so use this one for the figure and for reading comprehension, not for computation. It also carries a good follow-up the book poses directly: "if there had been a fourth type of volunteer, teenagers, what would df be?" Use for: new figure idea (O/E grid shaded by contribution) · quiz questions (~2, both about df)
p613 (printed 601) — Ex 88, college major × starting salary
What is there: A 5×3 table, N = 300, every cell a multiple of 5.
| Major | < $50k | $50–69k | $69k+ | Row total |
|---|---|---|---|---|
| English | 5 | 20 | 5 | 30 |
| Engineering | 10 | 30 | 60 | 100 |
| Nursing | 10 | 15 | 15 | 40 |
| Business | 10 | 20 | 30 | 60 |
| Psychology | 20 | 30 | 20 | 70 |
| Column total | 55 | 115 | 130 | 300 |
Why it improves the app: Round N and round marginals, df = 8, and a
scenario a psychology applicant will read with interest. Note the honest trap
built in: the English row (n = 30) has cells of 5, and the E ≥ 5 requirement is
right at the edge — usable as a follow-up question about the assumption.
Use for: worked problem · quiz questions (~3)
p605 (printed 593) — 11.3 Ex 23–25, "which test?" items, and Table 11.31
What is there: Three one-line scenarios asking only which test applies, then a 5×3 travel-distance × ticket-class table with N = 200 and printed marginals (row totals 41, 42, 48, 47, 22; column totals 73, 67, 60).
Why it improves the app: Unit 26's deliverable is literally "6 שאלות בחירת מבחן שבהן חי-בריבוע הוא אחד המסיחים", and Ex 24 and 25 are two ready-made ones where chi-square is the wrong answer for opposite reasons:
- Ex 24, player salary vs team winning percentage: both variables continuous → Pearson/regression, not χ².
- Ex 25, shoe brand vs run time: categorical IV, continuous DV → ANOVA, not χ².
- Ex 23, age group vs symptom presentation: two categoricals → χ² is right. That triple is exactly the discrimination step 2 of the app's decision tree tests. The Table 11.31 marginals give a clean extra df/expected-count drill (). Use for: quiz questions (~4, choosing-a-test) · worked problem
p615 (printed 603) — Ex 94 and 97, two true/false statements
What is there:
- Ex 94: "The number of degrees of freedom for a test of independence is equal to the sample size minus one." — false.
- Ex 97: "In a test of independence, the expected number is equal to the row total multiplied by the column total divided by the total surveyed." — true.
Why it improves the app: Two sentences, and they are the two things unit 26 says matter most. Translate almost verbatim into the "מה נכון / מה לא נובע" question style. Ex 94 in particular is the df error stated in its most tempting form. Use for: quiz questions (~2) · flashcards
p577–p579 (printed 565–567) — 11.2 Example 11.1, the E ≥ 5 requirement
What is there: The absenteeism example, whose only point is that the "12+" category has an expected count of 2, so it must be merged with "9–11" — and merging drops df from 4 to 3.
Why it improves the app: The app never states the E ≥ 5 assumption anywhere (OpenStax repeats it in a boxed NOTE on p577, p586 and p591). It belongs in the lesson as one line, and it makes a genuinely good exam question because the consequence is a change in df, not just a caveat: merging two categories both changes the χ² and reduces df. Modest but real. Use for: lesson prose (one paragraph) · quiz question (~1)
anova-oneway
The app's ANOVA example (SS 120/270/390, F = 6.00) is invented, equal-n, and significant. Chapter 13 supplies real datasets that break all three of those patterns.
p694–p695 (printed 682–683) — 13.2 Example 13.1, three diet plans
What is there: Ten weight-loss numbers in three unequally sized groups, plus the complete ANOVA table.
- Plan 1 (n = 4): 5, 4.5, 4, 3 — sum 16.5, mean 4.125
- Plan 2 (n = 3): 3.5, 7, 4.5 — sum 15, mean 5
- Plan 3 (n = 3): 8, 4, 3.5 — sum 15.5, mean 5.167
| Source | SS | df | MS | F |
|---|---|---|---|---|
| Between | 2.2458 | 2 | 1.1229 | 0.3769 |
| Within | 20.8542 | 7 | 2.9792 | |
| Total | 23.1 | 9 |
Why it improves the app: The best ANOVA find in the chapter, on three counts. (1) Unequal n. Every ANOVA the app currently shows is balanced, so and happen to agree, and the user can pass without ever learning which one is right. Here , , — and there is no group size to plug in. This is precisely the trap the lesson warns about ("דרגות החופש נגזרות ממספר הקבוצות ומ-N הכולל") but currently cannot demonstrate. (2) F < 1. The app asserts "F קרוב ל-1 מרמז על היעדר אפקט" but has never shown an F below 1. Here : the group means differ less than chance would predict. That is worth a paragraph. (3) I verified every number by hand — , , , — so the whole table is reconstructible from the ten raw numbers with no calculator. Use for: worked problem (full: from ten numbers to a complete table) · lesson prose (the F < 1 case) · quiz questions (~4, mostly df and "what does F = 0.38 mean")
p709 (printed 697) — 13.2 Ex 16–23, goals per game for four soccer teams
What is there: Four groups of five, every value a single digit 0–4.
- Team 1: 1, 2, 0, 3, 2 (sum 8)
- Team 2: 2, 3, 2, 4, 4 (sum 15)
- Team 3: 0, 1, 1, 0, 0 (sum 2)
- Team 4: 3, 4, 4, 3, 2 (sum 16)
Why it improves the app: Computed by hand: , , , , , df 3 / 16 / 19, , , , . This is the cleanest "build the whole ANOVA table from raw data" exercise in the chapter — twenty single-digit numbers producing a large, unambiguous F. And is exactly the shape of the app's favourite reverse question: from recover and . The exercise set (16–23) is already broken into the seven sub-answers the app's "fill in the missing cell" question type wants. Use for: worked problem · quiz questions (~5, one per table cell) · problem-bank ANOVA-completion drills
p701–p702 (printed 689–690) — 13.3 Example 13.4, bean plants
What is there: Three children, five plants each, integer heights — Tommy 24, 21, 23, 30, 23 · Tara 25, 31, 23, 20, 28 · Nick 23, 27, 22, 30, 20 — worked through a different route to F than sums of squares:
- group means 24.2, 25.4, 24.4; group variances 11.7, 18.3, 16.3
- variance of the three means = 0.413 →
- mean of the three variances = 15.433 = (pooled)
- , df 2 / 12
Why it improves the app: The lesson's central claim is that F is
signal ÷ noise, and it illustrates that with a figure
(anova-signal-noise.svg) — but the arithmetic it then shows goes through SS,
where the signal/noise reading is buried. This balanced-design shortcut is the
claim, literally: F = n × (spread of the group means) ÷ (average spread inside
the groups). Adding it gives the app a formula from which every row of its "מה
משנה את F" table can be read off in one line. Verified: all six intermediate
numbers reproduce exactly. It is also a second F < 1 case (means 24.2 / 25.4 /
24.4 against within-group variances near 15).
Use for: lesson prose (a short section, "F בדרך השנייה") · worked problem ·
new figure idea (three dot-columns whose means are annotated on one axis and
whose spreads are annotated on another, labelled as numerator and denominator)
p699–p700 (printed 687–688) — 13.3 Example 13.3, four sororities' grades
What is there: A balanced 4 × 5 design, N = 20, with the full ANOVA output: , , ; , , ; , p = 0.1241, tested at → do not reject. Raw data present (grade means 1.69–4.00).
Why it improves the app: Unit 22's headline point is that מובהקות and גודל אפקט are separate questions, and the app states it without an example. Here — the grouping explains 29.5% of the variance — and the result is still not significant, because N = 20. Pair it with the tomato example below (, significant) and you have a two-row table that makes the point permanently. The vs contrast is a second, independent teaching point: the same F would have been borderline at other α. Use for: lesson prose (η² vs significance, with numbers) · quiz questions (~3) · problem-bank item
p712 (printed 700) — 13.1 Ex 59, three traffic routes
What is there: Three groups of four, two-digit integers. Route 1: 30, 32, 27, 35 · Route 2: 27, 29, 28, 36 · Route 3: 16, 41, 22, 31.
Why it improves the app: Hand-computed: means 31, 30, 27.5 — visibly
different — yet , , , , F = 0.27.
Route 3 alone runs from 16 to 41. This is the most economical possible
demonstration of the app's own sentence " אינו מודד את גודל ההבדל בין
הממוצעים; הוא מודד אותו ביחס לרעש" — twelve numbers, and the means differ while
F says nothing is there. It is a better numeric companion to
anova-signal-noise.svg than anything currently in the file.
Use for: worked problem · lesson prose (annotating the signal/noise figure
with real numbers) · quiz question (~2)
p716–p717 (printed 704–705) — 13.3 Ex 72, Sanjay's paper airplanes
What is there: Three paper weights × four trials, and — unusually — the exercise is pre-structured into exactly the blanks of the Example 13.4 shortcut (variance of the group means → MS_between → mean of the sample variances → MS_within → F → df). Heavy: 5.1, 3.1, 4.7, 5.3 · Medium: 4.0, 3.5, 4.5, 6.1 · Light: 3.1, 3.3, 2.1, 1.9
Why it improves the app: The group means are 4.55, 4.525, 2.60 — two essentially identical and one clearly different. Computed: , , , df 2 / 9, significant at 5% but not at the 1% the exercise asks for. That single dataset carries three separate exam points at once: (a) a significant F while two of the three means are indistinguishable — the exact reason post-hoc tests exist, which is the "הטעות מס' 1" of unit 22; (b) the same F giving different decisions at α = .05 and α = .01; (c) the balanced-design shortcut, drilled. The pre-broken blank structure transfers directly into the app's worked-problem format. Use for: worked problem (the structured one) · quiz questions (~3, all "what may we conclude") · problem-bank item
p697–p698 (printed 685–686) — 13.3 Example 13.2, tomato mulch
What is there: A five-group, three-per-group design (N = 15) presented as a partially filled ANOVA table — SS and df given, MS and F left blank: (df 4), (df 10), (df 14); answer , p = 0.0248. Raw yields on p696 (2,625 … 9,230 grams).
Why it improves the app: The structure is the asset, not the numbers.
k = 5 with only 3 observations per group makes — a number you
cannot get from any single group size — and the printed table is already in
"fill in the blanks" form. , the significant partner to the
sorority example above. The gram-scale numbers are too large to compute by hand;
substitute Hebrew-appropriate quantities and keep the SS ratios.
Use for: ANOVA-completion drill (structure) · quiz questions (~2, df and η²)
p715–p716 (printed 703–704) — 13.3 Ex 69, statistics class delivery type
What is there: Final-exam scores by delivery mode with three different group sizes: Online 72, 84, 77, 80, 81 (n = 5); Hybrid 83, 73, 84, 81 (n = 4); Face-to-face 80, 78, 84, 81, 86, 79, 82 (n = 7).
Why it improves the app: N = 16, k = 3, — and the tempting wrong answer is nearly right, which makes it a better trap than a wildly wrong one. The scenario is also the app's own ANOVA example (שיטות הוראה) with real data behind it, so it drops straight into the existing lesson. Computed: , , — another honest null. Use for: quiz questions (~2, df) · problem-bank item
p716 (printed 704) — 13.3 Ex 70, meals eaten out, unequal groups
What is there: Four groups, single-digit counts, sizes 5 / 4 / 5 / 5 (N = 19): 6,8,2,4,6 · 4,1,5,2 · 7,3,5,4,6 · 8,3,5,1,7. Why it improves the app: Single digits and unequal n in one table — the easiest possible drill for . Cheap to import, useful as the warm-up before the diet-plan problem. The ethnic-group framing should be replaced with something neutral in Hebrew. Use for: quiz question (~1) · ANOVA-completion drill
regression
p650–p652 (printed 638–640) — 12.6 Example 12.12, the full residual table
What is there: The third-exam/final-exam dataset (n = 11) with every residual printed, then the consequence of deleting one point.
| x | y | ŷ | y − ŷ |
|---|---|---|---|
| 65 | 175 | 140 | +35 |
| 67 | 133 | 150 | −17 |
| 71 | 185 | 169 | +16 |
| 71 | 163 | 169 | −6 |
| 66 | 126 | 145 | −19 |
| 75 | 198 | 189 | +9 |
| 67 | 153 | 150 | +3 |
| 70 | 163 | 164 | −1 |
| 71 | 159 | 169 | −10 |
| 69 | 151 | 160 | −9 |
| 69 | 159 | 160 | −1 |
, , cutoff , so (65, 175) is flagged. Delete it and the line moves from , to , ; the prediction at x = 73 moves from 179.08 to 184.28.
Why it improves the app: The best regression find in the chapter, and it fills three separate gaps at once.
- The residual sum. The lesson asserts as an identity. Add these eleven residuals: 35 − 17 + 16 − 6 − 19 + 9 + 3 − 1 − 10 − 9 − 1 = 0, exactly. An asserted identity becomes a checkable one.
- A quantified outlier effect. The correlation lesson's manipulation table says "הוספה או הסרת חריג — תלוי במיקום". Here is what "תלוי" costs: one point out of eleven, and r goes 0.66 → 0.91 while the slope nearly doubles. The Anscombe section makes this point visually; this makes it numerically, which is what a "מה יקרה למתאם אם" question actually requires.
- with data behind it. The lesson gives with no worked instance; here is computed from SSE and then used, as the 2s outlier cutoff. The app already borrowed this dataset's derived numbers (the extrapolation example "טווח 65–75 מנבא 261 לתלמיד שקיבל 90" is exactly ), so importing the raw pairs is consistent with what is already there and lets the app compute rather than assert. Use for: worked problem (residual table + Σe = 0 + outlier removal) · lesson prose · new figure idea (the eleven points with the two fitted lines — before and after removing (65, 175) — overlaid, and the ±2s band drawn)
p655–p656 (printed 643–644) — Table 12.9, 95% critical values of r
What is there: A full table of the |r| needed for significance at α = .05 as a function of df = n − 2. Selected rows: df 1 → 0.997 · 2 → 0.950 · 3 → 0.878 · 4 → 0.811 · 5 → 0.754 · 6 → 0.707 · 7 → 0.666 · 8 → 0.632 · 9 → 0.602 · 10 → 0.576 · 12 → 0.532 · 15 → 0.482 · 17 → 0.456 · 20 → 0.423 · 30 → 0.349 · 40 → 0.304 · 50 → 0.273 · 60 → 0.250 · 80 → 0.217 · 100 → 0.195.
Why it improves the app: The correlation lesson teaches how to read r
(Cohen's benchmarks, R², what breaks it) but says nothing about how much
evidence a given r constitutes. This table is the missing axis: at n = 5 you
need r = 0.878 to clear α = .05; at n = 102 you need 0.195. That single sentence
connects unit 17 to unit 13 (עוצמה), to unit 36 (תוקף המסקנה הסטטיסטית) and to
the critique questions, where "the correlation was 0.7 but n = 6" and "the
correlation was 0.2 but n = 500" are both things a research description can say.
It is also a natural generated figure — a decaying curve of critical |r| against
n, which scripts/figures.mjs can compute rather than trace.
Use for: new figure idea (critical |r| vs n curve) · lesson prose ·
flashcard
p645–p646 (printed 633–634) — 12.4 Examples 12.7–12.10 and their Try Its
What is there: Nine short worked judgements pairing an r with an n and a verdict:
| r | n | critical value | significant? |
|---|---|---|---|
| −0.567 | 19 | 0.456 | yes |
| 0.708 | 9 | 0.666 | yes |
| 0.134 | 14 | 0.532 | no |
| 0.776 | 6 | 0.811 | no |
| −0.624 | 14 | 0.532 | yes |
| 0.6501 | 12 | 0.576 | yes |
| 0.5204 | 9 | 0.666 | no |
| −0.7204 | 8 | 0.707 | yes |
| 0.6631 | 11 | 0.602 | yes |
Why it improves the app: Rows 1 and 4 together are a first-class מתא"ם trap: r = 0.776 with n = 6 is not significant, while r = 0.567 with n = 19 is. The larger correlation is the weaker evidence. Every reflex the app has built so far ("עוצמה היא הערך המוחלט") points the wrong way here, which is exactly what makes it a good question. The set is nine ready-made items with the answers attached, including the degenerate r = 0 case (never significant at any n). Use for: quiz questions (~6) · lesson prose (one contrast) · flashcard
p648 (printed 636) — 12.6, outlier vs influential point
What is there: Two definitions kept explicitly apart. An outlier is far from the line vertically (a large residual). An influential point is far from the other points horizontally, in x, and "may have a big effect on the slope"; the test for it is to remove it and see whether the slope moves.
Why it improves the app: The app has no term for the second thing. Anscombe set 4 in the correlation lesson is an influential point ("כל זהים חוץ מאחד"), and the lesson describes it correctly, but without the name and without the general rule the reader cannot transfer it to a new scatter plot. Two sentences and a one-row table addition ("far in y = חריג; far in x = נקודה משפיעה") turn an anecdote into a rule. It also gives the app a sharper question type than "is this an outlier": which of these two flagged points changes the slope more, and why. Use for: lesson prose (correlation and regression both) · quiz questions (~2) · new figure idea (one scatter plot with a high-residual point and a high-leverage point marked differently, showing that only the second moves the line)
p677 (printed 665) — 12.6 Ex 74, coffee consumption and heart disease
What is there: Ten countries.
| Coffee (L/yr) | 2.5 | 3.9 | 2.9 | 2.4 | 2.9 | 0.8 | 9.1 | 2.7 | 0.8 | 0.7 |
|---|---|---|---|---|---|---|---|---|---|---|
| Heart-disease deaths | 221 | 167 | 131 | 191 | 220 | 297 | 71 | 172 | 211 | 300 |
The exercise asks which point has the largest residual and whether it is an outlier, an influential point, or both.
Why it improves the app: Nine of the ten countries sit between 0.7 and 3.9 litres; one sits at 9.1. That lone point is a textbook influential point — far in x, not far from the line — and it is doing most of the work in a negative relationship. And the scenario is a causation trap the app's own lesson would love: "more coffee, less heart disease" is exactly the "ככל ש… ⟹ בגלל ש…" inference the correlation lesson warns about, complete with an obvious third variable (national wealth). One dataset that serves range/leverage and causation critique, with n = 10 and no decimals worth fearing. Use for: worked problem · quiz questions (~3) · research-critique item
p635–p636 (printed 623–624) — 12.3 Example 12.6, the raw 11 (x, y) pairs
What is there: The dataset behind everything else in chapter 12 — third-exam score (out of 80) against final-exam score (out of 200):
| x | 65 | 67 | 71 | 71 | 66 | 75 | 67 | 70 | 71 | 69 | 69 |
|---|---|---|---|---|---|---|---|---|---|---|---|
| y | 175 | 133 | 185 | 163 | 126 | 198 | 153 | 163 | 159 | 151 | 159 |
, , , , , , and the line passes through as the app's lesson claims.
Why it improves the app: The app already uses this example's conclusions (the 65–75 range, the 261 extrapolation) without the data. Importing the eleven pairs closes the loop: the user can verify that the line goes through the means, that means 56% unexplained, and that the intercept −173.51 is the app's own "חותך בלי פירוש מעשי" made concrete — a third-exam score of 0 predicts a negative final grade, which is a better illustration than "משקל של אדם שגובהו 0". Also note the x range is only 65–75 while y runs 126–198, which is a free demonstration of the lesson's point that steepness ≠ strength. Use for: worked problem · lesson prose (the intercept illustration)
reading-results / choosing-a-test — smaller items
p694 (printed 682) — Table 13.1, the generic one-way ANOVA table
What is there: The blank ANOVA table with every cell given as a formula:
Factor row SS(Factor) | k−1 | SS(Factor)/(k−1) | MS(Factor)/MS(Error), Error
row SS(Error) | n−k | SS(Error)/(n−k), Total row SS(Total) | n−1.
Why it improves the app: The app's own table is equivalent, so this is not a
replacement — but the OpenStax naming (Factor / Error, alongside
Between / Within) is the second vocabulary a Hebrew exam question may use, and
the app currently teaches only one. Worth a synonym row in
choosing-a-test's "שמות נרדפים שכדאי לזהות" table, which already handles
ANOVA naming variants.
Use for: lesson prose (one table row) · flashcard
p692–p694 (printed 680–682) — 13.2 opening, and Figure 13.2
What is there: Two side-by-side sets of box plots — (a) H₀ true, three
identical distributions; (b) H₀ false, the same spreads shifted apart — with the
explanatory sentence "if the null hypothesis is false, then the variance of the
combined data is larger, which is caused by the different means."
Why it improves the app: That sentence is a cleaner statement of the variance
decomposition than "SS_T = SS_B + SS_W" alone: the total spread grows precisely
because the group means separated. The app's
anova-variance-decomposition.svg shows the algebra; a box-plot pair showing
combined-variance-grows would show the intuition. Idea only — the app must
generate its own.
Use for: new figure idea (two box-plot triples, H₀ true vs false, with the
pooled distribution drawn behind each)
Summary of what is actually worth doing first
If only five things get imported:
- p590 (578) — the anxiety × need-to-succeed contingency table. The app's biggest single content gap, in a psychology scenario, with round N.
- p694–695 (682–683) — the diet-plan ANOVA. Unequal n and F < 1, both absent from the app, both explicitly named traps in the spec.
- p650–652 (638–640) — the residual table and outlier removal. Turns three
asserted claims in
regression/lesson.mdinto checkable arithmetic. - p709 (697) — the soccer-goals ANOVA. Twenty single-digit numbers to a complete table; the ideal drill for the exam's most mechanical question type.
- p645–646 (633–634) + p655–656 (643–644) — r significance against n. An axis the app is missing entirely, and a good generated figure.