The video and action classification have extremely evolved by deep neural\nnetworks specially with two stream CNN using RGB and optical flow as inputs and\nthey present outstanding performance in terms of video analysis. One of the\nshortcoming of these methods is handling motion information extraction which is\ndone out side of the CNNs and relatively time consuming also on GPUs. So\nproposing end-to-end methods which are exploring to learn motion\nrepresentation, like 3D-CNN can achieve faster and accurate performance. We\npresent some novel deep CNNs using 3D architecture to model actions and motion\nrepresentation in an efficient way to be accurate and also as fast as\nreal-time. Our new networks learn distinctive models to combine deep motion\nfeatures into appearance model via learning optical flow features inside the\nnetwork.\n