Design it step by step
Define a worthwhile target
Invented task: predict which issue needs attention this week using only creation-time information. Later maintenance outcomes cannot enter the inputs. Use rules first where they already solve the task.
Make features interpretable
Noul can ask about reproduction steps; Score can rate a defined impact scale. Store feature IDs, questions, levels, model versions, and row results. Missing calls remain missing, not zero.
Split by entity or time
Keep updates and duplicates of one issue in one group. Train on training data, choose features on validation data, and reserve the final test. Feature proposers must not use test labels or errors.
Compare baselines before iterating
Use a majority-class, keyword, or simple text baseline. Hold budget and metrics fixed. Propose features from development errors and record retention decisions; many similar questions may add no information.
Run ablations and report scope
Remove features to test their contribution. Inspect class imbalance, important-issue misses, review burden, and cost. Freeze the design before final testing and report data time range, versions, and limits.
Worked example · Invented by this site
| Observation | Judgment | Application action |
|---|---|---|
| Feature says maintainer already marked urgent | Target or post-event leakage | Remove from prediction inputs |
| Updates of one issue split across sets | Near-duplicate leakage | Split by issue group |
| Validation improves after new questions | Generalization not yet established | Freeze and test the held-out set |
A copyable design draft
Original examples. JSON illustrates request or input structure; Python calculates invented scores without an API call. Verify current interfaces and task policy before integrating.
# Feature schema for an invented issue dataset.
{
"group_key": "issue_id",
"prediction_time": "issue_created_at",
"features": ["has_reproduction_steps", "impact_level"],
"missing_policy": "keep_missing_indicator",
"forbidden_inputs": ["future_maintainer_label", "closed_at"],
"split_roles": ["train", "validation", "untouched_test"]
}Design a decisionCommon mistakes
- Using test errors to propose new features.
- Treating missing as false.
- Reporting gains without baselines or costs.
TRY / THINK / COMPARE
Think first, then compare
Can you rewrite features after seeing final test errors and claim a new held-out gain on the same set?
Show explanation
That set is now development data. Disclose the reuse and obtain a new independent test or follow a predefined validation process.
Before handing it over
- Only prediction-time facts used.
- Entity and time leakage controlled.
- Baselines, versions, budget, and ablations retained.
- Final test did not guide features.
Common questions
Does automated question proposal complete research?
No. Data, training, validation, and budget control still matter; no training loop runs here.
Do features transfer to every project?
Not established. Revalidate across projects and languages.
Sources and further reading
Inspired by official patterns and Datawhale practice topics. Explanations, examples, and exercises are independently written. These are teaching designs, not live API runs or benchmarks.