Contextual Intelligence in Adaptive Wellness — Working Paper v2

Contextual
Intelligence in
Adaptive Wellness

When the right nudge arrives at the right moment: a conceptual framework, study protocol, and first longitudinal results on why and when SMS-based interventions change behavior in emerging adults.
Lorenzo Scardicchio  ·  Anjala Krishen  ·  Journey Research Group
Working Paper v2 — July 2026  ·  revised with the Fall ’25 → Spring ’26 longitudinal panel
“Would it be possible to do the full survey again? Beginning and end — and we can look at a lot of cool things at that point.”
Anjala Krishen, research planning session

I. The Problem

Most nudge research answers the wrong question

The dominant question in the digital wellbeing literature is: do nudges work? That question is too blunt. It treats nudges as a monolithic treatment, treats students as a monolithic population, and treats engagement as the outcome that matters. The result is research that can tell you average effects exist — usually small ones — while telling you almost nothing about why they exist, for whom, or under what conditions.

Journey’s university pilot, now spanning two semesters, produced a signal strong enough to suggest something more precise is happening. Students did not respond uniformly. Timing appeared to matter more than content — a pattern the new paired data now measures directly. Personalization functioned as a prerequisite, not a feature. Awareness increased before behavior changed. And the system’s scarcity constraint — only three nudges per week — forced every delivery decision to be a selection decision, not a scheduling decision.

That pattern raises a different question: what makes a nudge land? Not whether nudges work on average, but what specific combination of timing, context, personalization, and state-awareness turns a text message into a moment of genuine reflective interruption.

What most research asks

  • Did the intervention group improve more than control?
  • Is there a main effect of “nudging” on wellbeing?
  • Do students who receive messages report less screen time?
  • Is digital wellness messaging effective?

What this research asks

  • Why did some nudges land and others get ignored?
  • Does timing explain more variance than content?
  • Which students benefit and which disengage — and why?
  • Does the adaptive loop create lasting change or only short-term reaction?
“The interesting story to me here is how your initial state is affected. Do you actually create a permanent change? Or is it that we’re just working short term?”
— Anjala Krishen, research planning session

II. The Central Argument

Context is not a feature. Context is the mechanism.

Anind Dey defined context as “any information that can be used to characterize the situation of an entity.” In computing, context-awareness means using that information to provide relevant services. Journey extends that definition into behavioral intervention: context-awareness means using the student’s actual situation — their schedule, cognitive load, energy state, time of day, and likely emotional condition — to determine not just when to send a nudge, but what kind of nudge to send.

This is a stronger claim than “personalization helps.” It is a claim about mechanism. The hypothesis is that timing and context do more work than content in determining whether a nudge produces reflective interruption or gets dismissed. A perfectly worded nudge delivered at the wrong moment is noise. A simple question delivered at the moment of actual decision is an intervention.

The spring wave offers the first response-level support for this claim. Among the seven authorship items students rated in both waves, the two largest gains were “the nudges arrived at helpful times” (0% → 57% top-2 box) and “they made it easier to do what I wanted” (14% → 71%) — timing and friction, not message content. Section XIII presents the full grid.

Core Thesis
The Scarcity Constraint as Architectural Decision

Because Journey sends only three nudges per week, every delivery is a high-stakes selection decision. The question stops being “what should we say?” and becomes “when is the single best moment this week to say something this student can actually act on?”

Scarcity converts the timing engine from a scheduler into a selection engine. It forces the system to look for moments when an intervention is both possible (the student has a real break) and needed (the student’s likely state would benefit from a particular kind of support). The aim is not productivity. It is helping the student manage life holistically — across academics, stress, health, and connection.

Status: Observed in pilot

III. The Context Engine

Six dimensions evaluated before every nudge

Journey’s context engine draws on Dey’s four primary context categories — identity, location, activity, and time — and extends them with relational context and inferred internal state. The system does not send nudges and then hope they are relevant. It evaluates the student’s likely situation and selects the nudge type accordingly.

Fig. 1 — Context-Aware Nudge Selection
Schedule Time of Day Cognitive Load Student Profile Relational Context Internal State Selection Engine 3 nudges / week highest-leverage only NUDGE TYPE grounding / recovery / focus / social TONE hype / gentle / funny / deep TIMING pre-class / post / gap / evening CONTENT FRAME question / prompt / mirror

The Timing Logic

The system does not distribute nudges evenly across the week. It identifies specific windows where an intervention is both available (the student has a real break) and needed (the student’s likely state would benefit from a particular kind of support).

WindowPurposeLogic
Pre-ClassGrounding & IntentionBefore a class begins, a brief orienting prompt helps a student arrive mentally, not just physically. The nudge sets focus before the demand hits.
Post-ClassCapture & DecompressAfter class, the student either consolidates or dissipates. A reflection prompt catches learning while it is still warm.
After Stacked ClassesRecovery & RegulationAfter two or three back-to-back classes, the right nudge is not academic. It is physiological: stretch, walk, water, breathe. The system should know the difference.
Long GapsFocus Sprint or Social PromptA three-hour gap is an opportunity. Depending on the student’s goals and state, it could be a focus window or a connection window.

IV. The Tuesday Problem

Why Tuesday afternoon matters more than Monday morning

Consider a student with three consecutive classes on Tuesday, ending around 1:30 PM, followed by a gap before a 4 PM class. A generic system might send a focus reminder Monday morning because Monday feels like the “start” of the week. A contextually intelligent system recognizes that Tuesday afternoon — after cognitive depletion, before another demand — is the highest-leverage moment.

The nudge that lands there should not say “remember to study.” It should recognize that after three back-to-back classes, the student is likely drained. The most helpful intervention may be a recovery prompt: a walk, a stretch, a glass of water, a few minutes outside. That is not a concession. That is the system working correctly. It is selecting the intervention that matches the student’s likely nervous-system state, not their generic goal list.

Fig. 2 — A Student’s Tuesday: Where Context Selects the Nudge
Monday Tuesday Wednesday Thursday Friday 9:00 Psych 101 2:00 English 201 9:00 Bio Lab 10:30 Chem 102 12:00 Stats 200 ← NUDGE HERE 4:00 Seminar 9:00 Psych 101 2:00 English 201 10:30 Chem 102 12:00 Stats 200 9:00 Bio Lab
“After stacked classes (two to three back-to-back): purpose is recovery, reset nervous system. The best intervention after cognitive overload may be decompression, not more academic effort.”
— Journey Empathy Algorithm, design principle

Because the system only sends three nudges per week, it has to be highly selective. It should look for the moments when an intervention is both possible and needed: when a student has a real break, when stress is likely peaking, when energy is dipping, or when a small action could meaningfully improve focus, recovery, connection, or wellbeing. The aim is not just productivity, but helping the student manage life holistically.

V. The Adaptive Wellness Model

A feedback loop, not a broadcast system

Journey is not a messaging tool that happens to be personalized. It is an adaptive feedback system. The student enters an initial state. The system assesses that state. Nudges are delivered based on context. The student’s responses — and non-responses — feed back into the system. Over time, the system should learn and the student should grow. That adaptive loop is the novel research story.

Fig. 3 — The Adaptive Wellness Feedback Loop
STEP 01 Initial Assessment STEP 02 Context Engine STEP 03 SMS Delivery STEP 04 Response & Behavior STEP 05 Pulse Survey & Advisor feedback: steps 04–05 update steps 01–02
“How does adaptive wellness work — to me, that would be interesting. The paper would be really mapped to the software you’ve created.”
— Anjala Krishen

That adaptive loop maps directly onto Anjala’s framing of the research as feedback-control analysis. The initial survey establishes the baseline state. The context engine produces the intervention. The response data and follow-up surveys measure the output. The novel claim is that the system’s intelligence lies not in any single nudge, but in the loop itself — the ability to start from a student’s declared state, intervene at contextually selected moments, observe the response, and adjust. The fall→spring panel is the first full turn of that loop measured end to end.

VI. Illustrative Cases

Three student trajectories from the UNH panel, followed across the full year

Before aggregating, follow individual students through the year. Because the panel pairs each participant’s fall and spring responses, single trajectories can be traced end to end — and three of them, taken together, illustrate at the person level exactly the mechanisms the study is designed to test at the population level. One shows the mechanism working. One shows what happens when it fails. One complicates the durability story. All figures below are the students’ own paired survey responses; all quotes are verbatim.

Case 01 — S-01
The Responder Arc: Timing Converts Reading into Action
“Just by having time and motivation.” — S-01, fall wave, on what was holding them back

In fall, S-01 was a middling participant: read rate 4, action rate 3, satisfaction 3, “easier to act” 3, doom scrolling unchanged. Their stated obstacle was not knowledge or willingness — it was time and motivation, which is to say, a timing problem. By spring, every number had moved: read rate 7, action rate 6, satisfaction 6, “easier to act” 5, scrolling somewhat less.

The +3 action-rate gain is the largest in the panel. Nothing about this student’s goals changed between waves; what changed is that the nudges started meeting their actual windows. When the intervention found the moments where action was possible, the motivation problem largely dissolved. S-01 is the single-person version of H1: context-matched delivery does the work that generic encouragement cannot.

MeasureFallSpringΔ
Read rate (1–7)47+3
Action rate (1–7)36+3
Satisfaction (1–7)36+3
“Easier to act” (1–5)35+2
Doom scrollingNo changeSomewhat lessImproved
Case 02 — S-04
Personalization as Prerequisite: The Cost of One Generic Response
“When I told Journey I felt lonely cuz I had no friends it just told me to try joining clubs and I’m like 😑 cuz it was insensitive.” — S-04, fall wave

S-04 had the highest read rate in the fall panel — a 7. They read everything. And their fall satisfaction was a 2, one point off the floor, with “easier to act” at 2 and the one-word product verdict “make it more useful.” The combination is diagnostic: high read plus low satisfaction means the messages were reaching the student but not landing — and the open-text tells us exactly where the break happened. At a vulnerable moment, the student disclosed loneliness and received a generic suggestion. The content was not wrong; joining clubs is reasonable advice. It was contextually blind — the right information in the wrong register at the wrong moment — and it cost the system the student’s trust.

Spring shows a partial recovery: satisfaction doubled to 4, “easier to act” rose to 3, scrolling somewhat less. But the action rate never moved — 4 in fall, 4 in spring, the only both-substantive student with zero action gain. The trust that one contextually blind response spent was only partly won back over a full semester. S-04 is the single-person version of the paper’s central claim stated in the negative: when context fails, content cannot compensate.

MeasureFallSpringΔ
Read rate (1–7)75−2
Action rate (1–7)440
Satisfaction (1–7)24+2
“Easier to act” (1–5)23+1
Doom scrollingSomewhat lessSomewhat lessHeld
Case 03 — S-19
The Student Who Designed the Mechanism — and Then Regressed
“Having reminders of your goals throughout the day and then having prompts to reflect at night can help increase the amount of responses because people are more likely to be on their phones at night.” — S-19, fall wave

S-19 entered the study as its best case: fall read rate 6, action rate 5, satisfaction 6, and the only participant in either wave to report doom scrolling significantly less. And in the open-text, unprompted, this student proposed schedule-matched delivery and state-matched content — the paper’s thesis, reinvented independently by an eighteen-year-old describing their own phone habits. Struggling with “time management and setting aside time to do things for me,” they had already worked out that the answer was when, not what.

Then the complication: by spring, S-19’s read rate rose to 7 but their action rate fell to 3, and their scrolling reduction softened from “significantly less” to “somewhat less.” The panel’s strongest starter posted its largest decline. Three readings are possible — novelty decay, a ceiling effect (a student already near their best has nowhere to go but down), or a genuinely harder spring. The design cannot distinguish them. That is precisely why S-19 belongs in the paper: this case is Anjala’s short-term-reaction worry embodied, and it keeps the durability claim of H4 honest. The loop improved the panel on average; it did not improve everyone, and it did not hold its best case.

MeasureFallSpringΔ
Read rate (1–7)67+1
Action rate (1–7)53−2
“Easier to act” (1–5)44Held
Doom scrollingSignificantly lessSomewhat lessSoftened

These cases matter for the larger research because each one converts into a testable mechanism. S-01 tests whether timing-match converts reading into action (H1, RQ2). S-04 tests whether contextual blindness at vulnerable moments predicts persistent action-rate suppression — personalization as prerequisite, not feature (RQ4). S-19 tests whether early high engagers regress, and whether durability is a property of the loop or of the person (H4, RQ1). Each pattern visible at n=1 becomes a hypothesis testable across the full panel, and — with the next cohort — beyond it.

VII. Research Questions

Five questions that move the field

These questions are organized hierarchically. The first is the central question. The remaining four collectively build the explanatory model: not “prove the product works,” but “build a serious model of what is happening, for whom, why, and over what time horizon.”

RQ1
Lasting Change or Short-Term Reaction?
Does an adaptive, context-aware SMS intervention produce lasting changes in awareness, wellbeing, and self-directed behavior among emerging adults — or only short-term reactions?

This is the longitudinal question, and the panel now gives it a first descriptive answer. Across a full academic year, reported doom-scrolling reduction rose from 43% to 80% of substantive responders, satisfaction rose from 3.4 to 4.1 (1–7), and mean action rate rose from 3.1 to 3.8. Change accumulated over the year rather than decaying — though without a control group, maturation cannot be ruled out.

Status: First wave-pair measured — descriptive support
RQ2
Timing as Primary Mechanism
Does contextual timing — matching nudge delivery to the student’s schedule, cognitive load, and likely energy state — explain more variance in engagement and outcome than nudge content alone?

This tests the core claim of the contextual intelligence model. The spring grid gives it its first measured signal: “the nudges arrived at helpful times” posted the largest single gain of any item (0% → 57%), rising in lockstep with action rate. The definitive test still requires timing-matched vs. randomly timed comparison.

Status: Supported descriptively — causal test pending
RQ3
Engagement Segmentation
How do student engagement profiles — responders, avoiders, late responders, high-risk students — differ in baseline characteristics, response patterns, and longitudinal outcomes?

The panel’s pair types now map directly onto this taxonomy: 8 students substantive in both waves, 7 who moved from partial to substantive (the late responders Anjala predicted), and 4 who remained partial in both waves. Section X presents the observed segmentation.

Status: Segments observed in panel
RQ4
Proximal Mechanisms
Through what proximal mechanisms does the intervention operate — values alignment, mirror material, personalized tone, schedule-awareness, question-based framing, or scarcity — and can those mechanisms be distinguished from one another?

This is the question that prevents the paper from becoming a black-box evaluation. The authorship grid now orders the candidate mechanisms by measured movement — timing and friction-reduction first, personalization second, barrier-removal and tone third — but distinguishing them causally requires the MRT component.

Status: Mechanisms ordered — not yet separated
RQ5
Institutional Relevance
Do changes in awareness and values-behavior alignment predict downstream institutional outcomes — persistence, help-seeking, belonging, and academic self-efficacy?

This connects the awareness-first model to outcomes universities care about. It also tests the theoretical claim that awareness is upstream of behavior: if you increase a student’s capacity to notice and choose, downstream improvements follow. The survey panel cannot yet be joined to institutional records; this remains the next negotiation with the partner university.

Status: Theoretical

VIII. Study Design

Two studies, one explanatory model
Study 1
Baseline Landscape
What
Map the incoming student landscape before any intervention. Analyze initial survey data to identify correlations among goals, pain points, wellbeing, digital behavior, and demographic segments.
Protocol
Correlational analysis across intake variables. Cluster analysis to identify natural student segments. Path diagram showing baseline relationships among variables.
Measurement
Self-selected goals, stated challenges, wellbeing (ACIP / Healthy Minds Index), adapted QoL instrument, digital awareness items, demographic variables.
Output
A baseline model: what the world looks like before Journey does anything. No causal claims. The status quo, clearly mapped.
Study 2
Longitudinal + Mechanism Analysis
What
After a full year of adaptive nudging, test what changes, for whom, and through what mechanisms. The data collection for this study is now complete: same instrument, fall and spring, 19 paired participants.
Protocol
Comparable follow-up survey administered in spring. Engagement segmentation from pair types. Before-after comparison with paired change scores. Mechanism modeling: timing-match, personalization, and agency as mediators. Context features as predictors of individual nudge response.
Measurement
All Study 1 instruments repeated, plus: read rate and action rate (1–7 sliders), eight-item awareness grid, six-item helpfulness grid, seven-item authorship/agreement grid, doom-scrolling change, timing-vs-content forced choice, overall satisfaction, Sean Ellis “if Journey was gone” item, open-text voice of student.
Output
The core research contribution: does the adaptive loop produce durable change? For whom? Through what mechanisms? First results in Section XIII.

The Correlation vs. Causation Problem

Anjala’s most persistent concern is the directionality question: are students with better wellbeing simply more likely to respond to nudges, or do nudges actually improve wellbeing? This is not a minor methodological footnote. It is the central threat to the study’s credibility. A paper that accidentally implies causation from correlational data will fail with skeptical readers.

The design addresses this at multiple levels. The before-and-after structure allows person-level change scores — now computed for 8 both-substantive pairs. Engagement segmentation allows trajectory comparison across groups. Mechanism variables test whether specific features predict change. Future iterations will incorporate micro-randomized components enabling causal claims about specific features.

Epistemic Scope
What This Design Can and Cannot Prove

This design can establish what the baseline looks like, whether changes occur, whether changes differ by engagement profile, and whether mechanism variables predict change. It cannot definitively establish that Journey caused those changes. With a substantive base of 8 fall and 15 spring responders, within-person change must be reported as descriptive, not inferential.

The paper should state this clearly. The strength is in the explanatory model, not the causal claim. That is what makes it publishable: honest about limits, rich in mechanism.

Status: Design acknowledged

IX. Variables & Measurement

What gets measured, and how
VariableTypeInstrument / SourceWhen
Read rate / Action rateEngagementSurvey sliders (1–7, never→always)Fall + Spring
Digital Awareness (8 items)Outcome“Since using Journey, more aware of…” grid (1–5)Fall + Spring
Nudge Helpfulness (6 domains)OutcomeHelpfulness grid: mental health, focus, social, resources, physical, academic (1–5)Fall + Spring
Authorship / Agency (7 items)MediatorAgreement grid: personalized, goals, understood, timing, easier to act, barriers, not nagging (1–5)Fall + Spring
Doom ScrollingOutcomeSelf-report change item (5 options)Fall + Spring
Timing vs. ContentMechanismForced choice: “which was MORE important in making nudges effective?”Spring
Overall SatisfactionOutcome1–7 ratingFall + Spring
PMF SignalOutcomeSean Ellis: “how would you feel if Journey was gone?”Fall + Spring
Campus Resource UseOutcome“More likely to use campus resources” (1–7)Fall + Spring
Wellbeing (ACIP)OutcomeHealthy Minds Index (17 items)Planned: next cohort
Quality of LifeOutcomeAdapted college-student QoL (Krishen instrument)Planned: next cohort
Timing MatchMediatorSystem-generated: was nudge in high-leverage window?Per-nudge (logged)
Voice of StudentQualitativeOpen-text: struggles, “change one thing”Fall + Spring
Advisor NotesQualitativeStructured follow-up form (coded + open-ended)Per check-in

The measured instrument is narrower than the April protocol envisioned — the Healthy Minds Index and QoL scale were not fielded this year and move to the next cohort. What the fielded instrument gains is a direct mechanism probe: the forced-choice timing vs. content item asks students to adjudicate the paper’s central claim themselves.

“That first set of data would be interesting to analyze — what are the correlations? What kind of relationships are there? What do you see at an initial level in terms of a model?”
— Anjala Krishen, research planning session

X. Engagement Segmentation

Not one population — many. The segments are no longer hypothetical.

One of Anjala’s clearest analytical instincts is that the research gets much stronger when it stops treating students as one undifferentiated group. The paired panel now makes the segments observable: each student’s pattern of substantive vs. partial response across the two waves classifies them directly.

Segment A
Nudge Responders
Observed: 8 of 19 — substantive in both waves

Full answers fall and spring. Within this group, 5 of 8 increased their action rate over the year (mean Δ +0.9 on the 1–7 scale), 2 declined, 1 held flat. Doom-scrolling movement concentrated here: every both-substantive student who reported “no change” in fall reported “somewhat less” by spring except one.

Segment B
Late Responders
Observed: 7 of 19 — partial in fall, substantive in spring

The most interesting mechanism story, and it materialized: seven students who gave only uniform 4/4 read-action answers in fall completed the full instrument in spring — and 5 of 7 reported scrolling somewhat less. Did trust develop gradually? Did the nudges start landing? This group is the strongest argument for running interventions longer than one semester.

Segment C
Nudge Avoiders
Observed: 4 of 19 — partial in both waves

Enrolled but never substantively engaged with either survey wave. Not hostile — present enough to submit, disengaged enough to say nothing. Notification fatigue? Anti-AI sentiment? Wrong tone? This group is as important as the responders, and the fall open-text gives clues: “I would prefer it to be an app about wellness rather than texts.”

Segment D
High-Risk Students
Not yet identifiable in panel — survey cannot be joined to advisor flags

The de-identified survey carries no name or phone and cannot be linked to the advisor-flagged cohort. Testing whether the intervention helps those who need it most requires a consented join in the next cohort — a concrete protocol change this revision commits to.

The avoider segment is not a data loss. It is a research finding. Understanding why some students do not engage — and whether their reasons are addressable — prevents the study from reading as one-sided or promotional. It also generates the most interesting questions for intervention refinement.

“It would be interesting to map out types of people too — the nudge responders, the nudge avoiders, whatever.”
— Anjala Krishen

XI. Hypotheses

What we expected — and what the first paired wave says
H1
Context-Matched Nudges Outperform Random Timing
Nudges delivered in identified high-leverage windows (post-stacked-classes, pre-class, long-gap) will produce higher response rates and stronger proximal outcomes than nudges delivered at arbitrary times.

First evidence: “The nudges arrived at helpful times” moved from 0% to 57% top-2 box — the largest gain in the authorship grid — while action rate rose from 3.1 to 3.8. The pattern is consistent with H1 but is not the randomized comparison the claim ultimately requires.

Falsifiable if: timing-match scores do not predict engagement or outcomes.
Status: Descriptively supported
H2
Awareness Precedes Behavior Change
Students will report increased awareness of compulsive patterns before showing measurable reductions in problematic use. Awareness is upstream of behavior in the causal chain.

First evidence: the ordering is visible in the grain of the data. Awareness of mental-health patterns tripled (12.5% → 40%) and awareness of the need for breaks rose to 60%, while behavior moved to the moderate register: 12 of 15 spring responders report scrolling “somewhat less,” none “significantly less.” Awareness has moved further than behavior — exactly the sequence H2 predicts, caught mid-stride.

Falsifiable if: behavior change occurs without awareness change, or awareness change does not predict subsequent behavior change.
Status: Descriptively supported
H3
Engagement Profiles Predict Differential Outcomes
Responders, late responders, and avoiders will show different trajectories on wellbeing and awareness measures — and those trajectories will be partially predicted by baseline characteristics.

First evidence: the three profiles now exist and their spring outcomes differ (responders and late responders report scrolling reductions; avoiders report nothing at all). Predicting membership from baseline characteristics remains open — the fall wave of the late-responder group is, by definition, nearly empty.

Falsifiable if: segments do not differ on outcomes, or baseline variables do not predict segment membership.
Status: Segments observed — prediction untested
H4
The Adaptive Loop Produces Durable Change
Students who engaged with the adaptive system across the full year will show sustained improvement from baseline to follow-up, not merely spike-and-decay patterns.

First evidence: every headline metric improved from fall to spring rather than decaying — action rate, satisfaction, doom-scrolling reduction, and all seven authorship items. A spike-and-decay pattern would have produced the reverse. What the design cannot yet rule out is that the year itself, not the loop, matured the students.

Falsifiable if: improvements are concentrated in early weeks and decay, or post-intervention follow-up shows full regression to baseline.
Status: Consistent with durability

XII. Study Protocol

The concrete research plan — with progress marked
Phase 1: Baseline Analysis — Complete

Fall wave collected and analyzed. Correlations among goals, pain points, and engagement mapped. The fall picture: high read rates (mean 5.0), low action rates (3.1), personalization at zero, satisfaction at 3.4.

Phase 2: End-of-Year Survey — Complete

Spring wave fielded with the identical instrument before finals, as Anjala requested. 19 paired participants, 38 total submissions. This is the revision’s central new asset.

Phase 3: Engagement Segmentation — Complete (first pass)

Pair-type classification yields responders (8), late responders (7), and avoiders (4). High-risk identification blocked by de-identification — moves to next cohort with a consented join.

Phase 4: Longitudinal Analysis — In progress

Paired change scores computed descriptively (Section XIII). Remaining: mechanism modeling with timing-match as predictor, and the write-up that holds the descriptive line against the temptation to imply causation.

Phase 5: Qualitative Integration — Pending

Code open-text and advisor notes using structured categories. The fall criticisms (Section XIII) are the priority material — they explain the fall floor from which the spring gains rose.

Future
Micro-Randomized Trial (MRT) Component
What
Establish causal effects of specific nudge features by randomizing components (type, timing, tone, question vs. directive) within-person across occasions.
Proximal Outcomes
Time-to-open, session length, reply rate, self-reported urge/goal conflict, next-day digital behavior.
Moderators
Schedule density, time of day, baseline engagement profile, prior response pattern.
Why It Matters
The MRT converts the research from descriptive to genuinely causal at the feature level. The panel results sharpen its priority: the timing item moved most, so timing-matched vs. random delivery is the first randomization to run.

XIII. Evidence: The Fall → Spring Panel

What a year of adaptive nudging measurably changed — and what it honestly did not

The evidence base now has three layers, and the paper must keep them distinct. The platform layer: 74 first-year students enrolled, 14,222 total actions recorded, three nudges per week. The staff layer: the December 2025 implementation snapshot (89% engagement, 71% “felt personalized,” 57% “significantly less” doom scrolling) — aggregate, staff-reported. And now the response layer: 19 students who answered the same survey in fall and spring, paired by participant. Every figure below is measured from the export, not derived.

3.1 3.8
Action rate, substantive mean (1–7), fall → spring
43% 80%
Reported scrolling less, of those answering
14% 71%
“Nudges made it easier to act” (agree 4–5)
0% 50%
“Felt personalized to me” (agree 4–5)
3.4 4.1
Overall satisfaction, mean (1–7)
0% 57%
“Arrived at helpful times” — largest single gain
8 15
Substantive responders, fall → spring
3 / 14
“Very disappointed if Journey was gone” (spring) — 21%, below the 40% PMF benchmark

The Authorship Grid: Where the Movement Concentrated

The seven agreement items measure whether students experience the nudges as theirs — personalized, aligned with their goals, arriving at their moments. In fall, these items sat at or near the floor. By spring, every one had risen, and the ordering of the gains is the section’s central finding: the two largest gains are timing and friction, not message content.

Fig. 4 — Authorship Items, Top-2 Box, Fall vs Spring
Made it easier to do what I wanted 71% (+57) Arrived at helpful times 57% (+57) Reflected goals I actually care about 57% (+43) Felt personalized to me 50% (+50) Felt like Journey understood my life 50% (+50) Removed barriers to taking action 36% (+21) Felt like reminders, not nagging 36% (+7) Fall (n=7) Spring (n=14)

Awareness: The Upstream Variable Moves First

“Since using Journey, more aware of…”Fall (n=8)Spring (n=15)Δ
Your mental health patterns13%40%+27 pp
Opportunities to connect with others25%53%+28 pp
Your need for breaks during study sessions38%60%+23 pp
Career preparation steps13%27%+14 pp
Campus resources / personal goals38%47%+9 pp
When you’re feeling overwhelmed25%27%+2 pp
The importance of asking for help50%47%−3 pp

Note the exception: “asking for help” was already the fall ceiling item and did not move — consistent with the December staff observation that help-seeking recognition was the program’s earliest win. The items that moved most in spring are the ones that require accumulated self-observation: mental-health patterns, the need for breaks, opportunities to connect.

Doom Scrolling and the Honest Register

In spring, 12 of 15 responders report scrolling “somewhat less”; 3 report no change; none report “significantly less.” The behavior change is real, broad, and moderate. This matters because the December staff snapshot reported 57% scrolling “significantly less” — a figure the response-level data does not reproduce. The gap between staff impression and measured self-report is a finding, not an error, and this paper reports the measured number.

Finding
Relevance-Driven Engagement Without Addictive Design

The system was not optimized for engagement, yet substantive response nearly doubled from fall to spring. It did not tell students to stop scrolling, yet 80% of spring responders report scrolling less. It did not maximize session length, yet satisfaction rose and the only student who would have felt “relieved” if Journey disappeared in fall had no spring counterpart. That pattern — relevance-driven engagement without addictive design — now rests on measured, paired data rather than staff impression.

Status: Measured, descriptive

Voice of Student

The open-text answers, kept verbatim, explain the fall floor. They are the qualitative counterpart to the 0% fall personalization score:

“When I told Journey I felt lonely cuz I had no friends it just told me to try joining clubs and I’m like 😑 cuz it was insensitive”
— S-04, Fall wave. By spring, this student rated “easier to act” 3/5 and reported scrolling somewhat less.
“I think that having reminders of your goals throughout the day and then having prompts to reflect at night can help increase the amount of responses because people are more likely to be on their phones at night.”
— S-19, Fall wave. A student independently proposing timing-based delivery — the paper’s thesis, in the subject’s own words.
“I would prefer it to be an app about wellness rather than texts.”  ·  “More personalized responses.”  ·  “Increase the response speed to near immediate responses.”
— S-02, S-06, S-05, Fall wave. The criticisms cluster on medium, personalization, and latency — all addressable, all upstream of the spring gains.

Limits of the Panel

Four constraints bound every claim above. The substantive base is small — 8 fall, 15 spring, 8 both-substantive pairs — so change is reported descriptively, never inferentially. Wave labeling follows export chronology and must be confirmed against source timestamps before publication. Five students submitted uniform 4/4 partials in one or both waves and are excluded from substantive figures. And the de-identified panel cannot be joined to the named cohort, which blocks the high-risk analysis until next year’s consented design. Stating these limits plainly is not a weakness of the paper; it is Contribution 04.

XIV. Theoretical Framework

Where this sits in the literature

The paper integrates five theoretical streams. None is new individually. The contribution is in showing how they interact within a single adaptive system and testing that interaction empirically.

StreamCore IdeaHow Journey Extends It
Context-Aware Computing
(Dey, 2001)
Any information that characterizes the situation of an entity — and its use to provide relevant services.Extends the definition from computing into behavioral intervention.
ACIP Framework
(Davidson et al., 2020)
Awareness, Connection, Insight, and Purpose as trainable dimensions of wellbeing.The awareness grid tracks the A of ACIP directly; the Healthy Minds Index enters with the next cohort.
Just-in-Time Adaptive Interventions
(Nahum-Shani et al., 2018)
Delivering the right support at the right time by adapting to changing context and state.Implements JITAI logic through schedule-awareness and scarcity-constrained selection — and now shows the timing item moving most.
Motivational Interviewing
(Miller & Rollnick, 2013)
Resolving ambivalence by evoking the person’s own reasons for change.Question-first, reflective, permission-giving nudge architecture.
Digital Harm ReductionAn awareness-first frame replacing abstinence logic with agency-building.The target is not less screen time but more intentional technology use — mirrored in the “somewhat less, not significantly less” spring register.

XV. The Contribution

Why this paper advances the field
Contribution 01
Context as Mechanism, Not Decoration

Most nudge studies treat timing as a delivery parameter. This paper treats it as the primary mechanism of action — and its first paired wave found the timing item posting the largest measured gain in the instrument.

Contribution 02
Scarcity-Constrained Selection

The three-nudges-per-week constraint introduces a novel design logic: intervention value comes not from frequency but from precision.

Contribution 03
Adaptive Loop as Research Object

The paper does not treat the system as a black box. It maps the adaptive loop explicitly, and the fall→spring panel constitutes the loop’s first measured full rotation: state in, intervention, state out.

Contribution 04
Honest About Causal Limits

Shows the baseline clearly, measures change, models mechanisms, accounts for segmentation — and corrects its own earlier evidence, replacing staff-reported figures with lower, measured ones and calling the gap a finding. Scientific honesty as contribution.

“The second study could be looking at how the nudges sort of work with what we find in the first study.”
— Anjala Krishen

A paper that says “nudges helped students” is common. A paper that says “here is the baseline model, here is the adaptive system, here is the measured within-person change, here are the plausible mechanisms ordered by movement, here is who it works for and who it does not, and here are the causal limits” — that is a contribution. The discussion should maintain a healthy scientific tension: optimistic enough to show promise, sober enough to acknowledge what the design cannot prove.

The deepest question the discussion should address is whether adaptive, context-aware wellness interventions represent a genuinely different category of digital intervention — one built around awareness and agency rather than restriction and compliance — and whether that difference matters for how emerging adults develop the capacity to direct their own cognitive and emotional lives. The panel’s answer, so far, is a cautious yes: the students moved, in their own words, from a system that “just told me to try joining clubs” to one whose nudges “arrived at helpful times.”