> For the complete documentation index, see [llms.txt](https://lwang010.gitbook.io/longw/llms.txt). Markdown versions of documentation pages are available by appending `.md` to page URLs; this page is available as [Markdown](https://lwang010.gitbook.io/longw/mlops/chap-5.-lessons-learnt-from-paper-reproduction-1/case-study.md).

# 5.1 Case study

The common protocol for reproducing the papers includes:

* paper with code available
  * clone the git repo and create a container with the same environment (ubuntu/python/tf/pytorch etc) as the README suggested
  * download the dataset they used for training and create a mini dataset for debugging
  * run on the whole training pipeline to make sure it is reproducible (compare the loss/metrics)
  * modified the data loader/loss func/#channel etc to make it compatible with the customized data we are working at
  * get a baseline
  * run training and record the experiments
* paper without code available
  * chose the paper in the fields you're familiar
  * build an MVP first (don't investigate too much time before you're convinced by the performance of the MVP)
  * change a module once a time
