The following is a brief directory structure and description for this example:
├── data # Data set directory
│ ├── prepare_data.sh # Shell script to download and process dataset
│ └── README.md # Documentation describing how to prepare dataset
│ └── script # Directory contains scripts to process dataset
│ ├── generate_neg.py # Create negative sample
│ ├── generate_voc.py # Create a list of features
│ ├── history_behavior_list.py # Count user's history behaviors
│ ├── item_map.py # Create a map between item id and cate
│ ├── local_aggretor.py # Generate sample data
│ ├── pick2txt.py # Convert voc's format
│ ├── process_data.py # Parse raw json data
│ └── split_by_user.py # Divide the dataset
├── distribute_k8s # Distributed training related files
│ ├── distribute_k8s_BF16.yaml # k8s yaml to crate a training job with BF16 feature
│ ├── distribute_k8s_FP32.yaml # k8s yaml to crate a training job
│ └── launch.py # Script to set env for distributed training
├── README.md # Documentation
├── result # Output directory
│ └── README.md # Documentation describing output directory
└── train.py # Training script
Deep Interest Evolution Network(DIEN) is proposed by Alibaba in 2018.11 which is a click-through rate(CRT) prediction model for e-commerce industry, focusing on capturing temporal interests from the user's historical behavior sequence.
-
Please prepare the data set and DeepRec env.
- Manually
- Follow dataset preparation to prepare data set.
- Download code by
git clone https://github.com/alibaba/DeepRec - Follow How to Build to build DeepRec whl package and install by
pip install $DEEPREC_WHL.
- Docker(Recommended)
docker pull alideeprec/deeprec-release-modelzoo:latest docker run -it alideeprec/deeprec-release-modelzoo:latest /bin/bash # In docker container cd /root/modelzoo/dien
- Manually
-
Training.
python train.py # Memory acceleration with jemalloc. # The required ENV `MALLOC_CONF` is already set in the code. LD_PRELOAD=./libjemalloc.so.2.5.1 python train.pyUse argument
--bf16to enable DeepRec BF16 feature.python train.py --bf16 # Memory acceleration with jemalloc. # The required ENV `MALLOC_CONF` is already set in the code. LD_PRELOAD=./libjemalloc.so.2.5.1 python train.py --bf16In the community tensorflow environment, use argument
--tfto disable all of DeepRec's feature.python train.py --tfUse arguments to set up a custom configuation:
- DeepRec Features:
export START_STATISTIC_STEPandexport STOP_STATISTIC_STEP: Set ENV to configure CPU memory optimization. This is already set to 100 & 110 in the code by default.--bf16: Enable DeepRec BF16 feature in DeepRec. Use FP32 by default.--emb_fusion: Whether to enable embedding fusion, Default to True.--op_fusion: Whether to enable Auto graph fusion feature. Default to True.--optimizer: Choose the optimizer for deep model from ['adam', 'adamasync', 'adagraddecay']. Use adagrad by default.--smartstaged: Whether to enable smart staged feature of DeepRec, Default to True.--micro_batch: Set num for Auto Mirco Batch. Default 0 to close.(Not really enabled)--ev: Whether to enable DeepRec EmbeddingVariable. Default to False.--group_embedding: Use GroupEmbedding features.--adaptive_emb: Whether to enable Adaptive Embedding. Default to False.--ev_elimination: Set Feature Elimination of EmbeddingVariable Feature. Options [None, 'l2', 'gstep'], default to None.--ev_filter: Set Feature Filter of EmbeddingVariable Feature. Options [None, 'counter', 'cbf'], default to None.--dynamic_ev: Whether to enable Dynamic-dimension Embedding Variable. Default to False.(Not really enabled)--multihash: Whether to enable Multi-Hash Variable. Default to False.(Not really enabled)--incremental_ckpt: Set time of save Incremental Checkpoint. Default 0 to close.--workqueue: Whether to enable Work Queue. Default to False.--protocol: Set the protocol ['grpc', 'grpc++', 'star_server'] used when starting server in distributed training. Default to grpc.--parquet_dataset: Whether to enable ParquetDataset. Default isTrue.--parquet_dataset_shuffle: Whether to enable shuffle operation for Parquet Dataset. Default toFalse.
- Basic Settings:
--data_location: Full path of train & eval data, default to./data.--steps: Set the number of steps on train dataset. Default will be set to 1 epoch.--no_eval: Do not evaluate trained model by eval dataset.--batch_size: Batch size to train. Default to 2048.--output_dir: Full path to output directory for logs and saved model, default to./result.--checkpoint: Full path to checkpoints input/output directory, default to$(OUTPUT_DIR)/model_$(MODEL_NAME)_$(TIMESTAMPS)--save_steps: Set the number of steps on saving checkpoints, zero to close. Default will be set to 0.--seed: Set the random seed for tensorflow.--timeline: Save steps of profile hooks to record timeline, zero to close, defualt to 0.--keep_checkpoint_max: Maximum number of recent checkpoint to keep. Default to 1.--learning_rate: Learning rate for model. Default to 0.001.--inter: Set inter op parallelism threads. Default to 0.--intra: Set intra op parallelism threads. Default to 0.--input_layer_partitioner: Slice size of input layer partitioner(units MB).--dense_layer_partitioner: Slice size of dense layer partitioner(units kB).--tf: Use TF 1.15.5 API and disable DeepRec features.
- DeepRec Features:
- Prepare a K8S cluster. Alibaba Cloud ACK Service(Alibaba Cloud Container Service for Kubernetes) can quickly create a Kubernetes cluster.
- Perpare a shared storage volume. For Alibaba Cloud ACK, OSS(Object Storage Service) can be used as a shared storage volume.
- Create a PVC(PeritetVolumeClaim) named
deeprecfor storage volumn in cluster. - Prepare docker image.
alideeprec/deeprec-release-modelzoo:latestis recommended. - Create a k8s job from
.yamlto run distributed training.kubectl create -f $YAML_FILE - Show training log by
kubectl logs -f trainer-worker-0
The benchmark is performed on the Alibaba Cloud ECS general purpose instance family with high clock speeds - ecs.g8i.4xlarge.
-
Hardware
- Model name: Intel(R) Xeon(R) Platinum 8475B
- CPU(s): 16
- Socket(s): 1
- Core(s) per socket: 8
- Thread(s) per core: 2
- Memory: 64G
-
Software
- kernel: Linux version 5.15.0-58-generic (buildd@lcy02-amd64-101)(AMX patched)
- OS: Ubuntu 22.04.2 LTS
- GCC: 11.3.0
- Docker: 20.10.21
| Framework | DType | Accuracy | AUC | Throughput | |
| DIEN | Community TensorFlow | FP32 | 0.575529 | 0.597272 | 6327.50(baseline) |
| DeepRec w/ oneDNN | FP32 | 0.543935 | 0.5972728 | 10094.21(1.60x) | |
| DeepRec w/ oneDNN | FP32+BF16 | 0.551233 | 0.597272 | 11565.63(1.83x) |
- Community TensorFlow version is v1.15.5.
The benchmark is performed on the Alibaba Cloud ACK Service(Alibaba Cloud Container Service for Kubernetes), the K8S cluster is composed of the following ten machines.
- Hardware
- Model name: Intel(R) Xeon(R) Platinum 8369HC CPU @ 3.30GHz
- CPU(s): 8
- Socket(s): 1
- Core(s) per socket: 4
- Thread(s) per core: 2
- Memory: 32G
| Framework | Protocol | DType | Throughput | |
| DIEN | Community TensorFlow | GRPC | FP32 | |
| DeepRec w/ oneDNN | GRPC | FP32 | ||
| DeepRec w/ oneDNN | GRPC | FP32+BF16 |
- Community TensorFlow version is v1.15.5.
Amazon Dataset Books dataset is used as benchmark dataset.
We provide the dataset in two formats:
- CSV Format
Put data file into ./data/
For details of Data download, see Data Preparation. We also provide these files at Amazon CSV Dataset(Recommended). - Parquet Format
Put data file into ./data/ These files are available at Amazon Parquet Dataset.
- cat_voc.pkl: Contain a list of book categories.
- mid_voc.pkl: Contain a list of item id(book id).
- uid_voc.pkl: Contain a list of user id.
- reviews-info: Contain a list of user's review.
Each piece of data is as:<user id> <item id> <rating score> <timestamp> - item-info: Contain mapping relationship between item id and categories.
<item id> <categories> - local_train_splitByUser & local_test_splitByUser: Train and evaluate dataset, consist of user id, item info, user's historical behavior.
Each piece of data is as:<label> <user id> <item id> <categories> <history item id list> <history item categories list>
The history data are splitted by'�'
Reviews are regard as behaviors and those from one user are sort by time. Assuming user u has T behaviors, the first T-1 behaviors are used to predict whether user u will write the T-th review.