In dialog studies, we often encode a dialog using a hierarchical encoder\nwhere each utterance is converted into an utterance vector, and then a sequence\nof utterance vectors is converted into a dialog vector. Since knowing who\nproduced which utterance is essential to understanding a dialog, conventional\nmethods tried integrating speaker labels into utterance vectors. We found the\nmethod problematic in some cases where speaker annotations are inconsistent\namong different dialogs. A relative speaker modeling method is proposed to\naddress the problem. Experimental evaluations on dialog act recognition and\nresponse generation show that the proposed method yields superior and more\nconsistent performances.\n