Deep convolutional neural networks (CNNs) have shown a strong ability in\nmining discriminative object pose and parts information for image recognition.\nFor fine-grained recognition, context-aware rich feature representation of\nobject/scene plays a key role since it exhibits a significant variance in the\nsame subcategory and subtle variance among different subcategories. Finding the\nsubtle variance that fully characterizes the object/scene is not\nstraightforward. To address this, we propose a novel context-aware attentional\npooling (CAP) that effectively captures subtle changes via sub-pixel gradients,\nand learns to attend informative integral regions and their importance in\ndiscriminating different subcategories without requiring the bounding-box\nand/or distinguishable part annotations. We also introduce a novel feature\nencoding by considering the intrinsic consistency between the informativeness\nof the integral regions and their spatial structures to capture the semantic\ncorrelation among them. Our approach is simple yet extremely effective and can\nbe easily applied on top of a standard classification backbone network. We\nevaluate our approach using six state-of-the-art (SotA) backbone networks and\neight benchmark datasets. Our method significantly outperforms the SotA\napproaches on six datasets and is very competitive with the remaining two.\n