<?xml version="1.0" encoding="ISO-8859-1"?><article xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink" xmlns:xsi="http://www.w3.org/2001/XMLSchema-instance">
<front>
<journal-meta>
<journal-id>1405-5546</journal-id>
<journal-title><![CDATA[Computación y Sistemas]]></journal-title>
<abbrev-journal-title><![CDATA[Comp. y Sist.]]></abbrev-journal-title>
<issn>1405-5546</issn>
<publisher>
<publisher-name><![CDATA[Instituto Politécnico Nacional, Centro de Investigación en Computación]]></publisher-name>
</publisher>
</journal-meta>
<article-meta>
<article-id>S1405-55462018000401223</article-id>
<article-id pub-id-type="doi">10.13053/cys-22-4-3009</article-id>
<title-group>
<article-title xml:lang="en"><![CDATA[Tunisian Dialect Sentiment Analysis: A Natural Language Processing-based Approach]]></article-title>
</title-group>
<contrib-group>
<contrib contrib-type="author">
<name>
<surname><![CDATA[Mulki]]></surname>
<given-names><![CDATA[Hala]]></given-names>
</name>
<xref ref-type="aff" rid="Aff"/>
</contrib>
<contrib contrib-type="author">
<name>
<surname><![CDATA[Haddad]]></surname>
<given-names><![CDATA[Hatem]]></given-names>
</name>
<xref ref-type="aff" rid="Aff"/>
</contrib>
<contrib contrib-type="author">
<name>
<surname><![CDATA[Ali]]></surname>
<given-names><![CDATA[Chedi Bechikh]]></given-names>
</name>
<xref ref-type="aff" rid="Aff"/>
</contrib>
<contrib contrib-type="author">
<name>
<surname><![CDATA[Babao&#287;lu]]></surname>
<given-names><![CDATA[Ismail]]></given-names>
</name>
<xref ref-type="aff" rid="Aff"/>
</contrib>
</contrib-group>
<aff id="Af1">
<institution><![CDATA[,Selcuk University Department of Computer Engineering ]]></institution>
<addr-line><![CDATA[ ]]></addr-line>
<country>Turkey</country>
</aff>
<aff id="Af2">
<institution><![CDATA[,Université Libre de Bruxelles Department of Computer &amp; Decision Engineering ]]></institution>
<addr-line><![CDATA[ ]]></addr-line>
<country>Belgium</country>
</aff>
<aff id="Af3">
<institution><![CDATA[,Carthage University LISI Laboratory ]]></institution>
<addr-line><![CDATA[ ]]></addr-line>
<country>Tunisia</country>
</aff>
<pub-date pub-type="pub">
<day>00</day>
<month>12</month>
<year>2018</year>
</pub-date>
<pub-date pub-type="epub">
<day>00</day>
<month>12</month>
<year>2018</year>
</pub-date>
<volume>22</volume>
<numero>4</numero>
<fpage>1223</fpage>
<lpage>1232</lpage>
<copyright-statement/>
<copyright-year/>
<self-uri xlink:href="http://www.scielo.org.mx/scielo.php?script=sci_arttext&amp;pid=S1405-55462018000401223&amp;lng=en&amp;nrm=iso"></self-uri><self-uri xlink:href="http://www.scielo.org.mx/scielo.php?script=sci_abstract&amp;pid=S1405-55462018000401223&amp;lng=en&amp;nrm=iso"></self-uri><self-uri xlink:href="http://www.scielo.org.mx/scielo.php?script=sci_pdf&amp;pid=S1405-55462018000401223&amp;lng=en&amp;nrm=iso"></self-uri><abstract abstract-type="short" xml:lang="en"><p><![CDATA[Abstract: Social media platforms have been witnessing a significant increase in posts written in the Tunisian dialect since the uprising in Tunisia at the end of 2010. Most of the posted tweets or comments reflect the impressions of the Tunisian public towards social, economical and political major events. These opinions have been tracked, analyzed and evaluated through sentiment analysis systems. In the current study, we investigate the impact of several preprocessing techniques on sentiment analysis using two sentiment classification models: Supervised and lexicon-based. These models were trained on three Tunisian datasets of different sizes and multiple domains. Our results emphasize the positive impact of preprocessing phase on the evaluation measures of both sentiment classifiers as the baseline was significantly outperformed when stemming, emoji recognition and negation detection tasks were applied. Moreover, integrating named entities with these tasks enhanced the lexicon-based classification performance in all datasets and that of the supervised model in medium and small sized datasets.]]></p></abstract>
<kwd-group>
<kwd lng="en"><![CDATA[Tunisian sentiment analysis]]></kwd>
<kwd lng="en"><![CDATA[text preprocessing]]></kwd>
<kwd lng="en"><![CDATA[named entities]]></kwd>
</kwd-group>
</article-meta>
</front><back>
<ref-list>
<ref id="B1">
<label>1</label><nlm-citation citation-type="confpro">
<person-group person-group-type="author">
<name>
<surname><![CDATA[Abdelali]]></surname>
<given-names><![CDATA[A.]]></given-names>
</name>
<name>
<surname><![CDATA[Darwish]]></surname>
<given-names><![CDATA[K.]]></given-names>
</name>
<name>
<surname><![CDATA[Durrani]]></surname>
<given-names><![CDATA[N.]]></given-names>
</name>
<name>
<surname><![CDATA[Mubarak]]></surname>
<given-names><![CDATA[H.]]></given-names>
</name>
</person-group>
<source><![CDATA[Farasa: A fast and furious segmenter for arabic]]></source>
<year>2016</year>
<conf-name><![CDATA[ 2016 Conference of the North American Chapter of the Association for Computational Linguistics: Demonstrations]]></conf-name>
<conf-loc> </conf-loc>
<page-range>11-6</page-range></nlm-citation>
</ref>
<ref id="B2">
<label>2</label><nlm-citation citation-type="confpro">
<person-group person-group-type="author">
<name>
<surname><![CDATA[Akaichi]]></surname>
<given-names><![CDATA[J]]></given-names>
</name>
</person-group>
<source><![CDATA[Sentiment classification at the time of the tunisian uprising: Machine learning techniques applied to a new corpus for arabic language]]></source>
<year>2014</year>
<conf-name><![CDATA[ Network Intelligence Conference (ENIC), 2014 European]]></conf-name>
<conf-loc> </conf-loc>
<page-range>38-45</page-range></nlm-citation>
</ref>
<ref id="B3">
<label>3</label><nlm-citation citation-type="confpro">
<person-group person-group-type="author">
<name>
<surname><![CDATA[Aly]]></surname>
<given-names><![CDATA[M.]]></given-names>
</name>
<name>
<surname><![CDATA[Atiya]]></surname>
<given-names><![CDATA[A.]]></given-names>
</name>
</person-group>
<source><![CDATA[Labr: A large scale arabic book reviews dataset]]></source>
<year>2013</year>
<volume>2</volume>
<conf-name><![CDATA[ 51st Annual Meeting of the Association for Computational Linguistics]]></conf-name>
<conf-loc> </conf-loc>
<page-range>494-8</page-range></nlm-citation>
</ref>
<ref id="B4">
<label>4</label><nlm-citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname><![CDATA[Brahimi]]></surname>
<given-names><![CDATA[B.]]></given-names>
</name>
<name>
<surname><![CDATA[Touahria]]></surname>
<given-names><![CDATA[M.]]></given-names>
</name>
<name>
<surname><![CDATA[Tari]]></surname>
<given-names><![CDATA[A.]]></given-names>
</name>
</person-group>
<article-title xml:lang=""><![CDATA[Data and text mining techniques for classifying arabic tweet polarity]]></article-title>
<source><![CDATA[Journal of Digital Information Management]]></source>
<year>2016</year>
<volume>14</volume>
<numero>1</numero>
<issue>1</issue>
</nlm-citation>
</ref>
<ref id="B5">
<label>5</label><nlm-citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname><![CDATA[Duwairi]]></surname>
<given-names><![CDATA[R.]]></given-names>
</name>
<name>
<surname><![CDATA[El-Orfali]]></surname>
<given-names><![CDATA[M.]]></given-names>
</name>
</person-group>
<article-title xml:lang=""><![CDATA[A study of the effects of preprocessing strategies on sentiment analysis for arabic text]]></article-title>
<source><![CDATA[Journal of Information Science]]></source>
<year>2014</year>
<volume>40</volume>
<numero>4</numero>
<issue>4</issue>
<page-range>501-13</page-range></nlm-citation>
</ref>
<ref id="B6">
<label>6</label><nlm-citation citation-type="confpro">
<person-group person-group-type="author">
<name>
<surname><![CDATA[El-Beltagy]]></surname>
<given-names><![CDATA[S. R.]]></given-names>
</name>
<name>
<surname><![CDATA[El kalamawy]]></surname>
<given-names><![CDATA[M.]]></given-names>
</name>
<name>
<surname><![CDATA[Soliman]]></surname>
<given-names><![CDATA[A. B.]]></given-names>
</name>
</person-group>
<source><![CDATA[Niletmrg at semeval-2017 task 4: Arabic sentiment analysis]]></source>
<year>2017</year>
<conf-name><![CDATA[ 11th International Workshop on Semantic Evaluation (SemEval-2017)]]></conf-name>
<conf-loc> </conf-loc>
<page-range>790-5</page-range></nlm-citation>
</ref>
<ref id="B7">
<label>7</label><nlm-citation citation-type="confpro">
<person-group person-group-type="author">
<name>
<surname><![CDATA[Gridach]]></surname>
<given-names><![CDATA[M]]></given-names>
</name>
</person-group>
<source><![CDATA[Character-aware neural net-works for arabic named entity recognition for social media]]></source>
<year>2016</year>
<conf-name><![CDATA[ 6th Workshop on South and Southeast Asian Natural Language Processing (WSSANLP2016)]]></conf-name>
<conf-loc> </conf-loc>
<page-range>23-32</page-range></nlm-citation>
</ref>
<ref id="B8">
<label>8</label><nlm-citation citation-type="book">
<person-group person-group-type="author">
<name>
<surname><![CDATA[Karmani]]></surname>
<given-names><![CDATA[N]]></given-names>
</name>
</person-group>
<source><![CDATA[Tunisian Arabic Customer&#8217;s Reviews Processing and Analysis for an Internet Supervision System]]></source>
<year>2017</year>
<publisher-loc><![CDATA[Tunisia ]]></publisher-loc>
<publisher-name><![CDATA[Sfax University]]></publisher-name>
</nlm-citation>
</ref>
<ref id="B9">
<label>9</label><nlm-citation citation-type="confpro">
<person-group person-group-type="author">
<name>
<surname><![CDATA[Larkey]]></surname>
<given-names><![CDATA[L. S.]]></given-names>
</name>
<name>
<surname><![CDATA[Ballesteros]]></surname>
<given-names><![CDATA[L.]]></given-names>
</name>
<name>
<surname><![CDATA[Connell]]></surname>
<given-names><![CDATA[M. E.]]></given-names>
</name>
</person-group>
<source><![CDATA[Improving stemming for arabic information retrieval: light stemming and co-occurrence analysis]]></source>
<year>2002</year>
<conf-name><![CDATA[ 25th annual international ACM SIGIR conference on Research and development in information retrieval]]></conf-name>
<conf-loc> </conf-loc>
<page-range>275-82</page-range></nlm-citation>
</ref>
<ref id="B10">
<label>10</label><nlm-citation citation-type="confpro">
<person-group person-group-type="author">
<name>
<surname><![CDATA[Malmasi]]></surname>
<given-names><![CDATA[S.]]></given-names>
</name>
<name>
<surname><![CDATA[Refaee]]></surname>
<given-names><![CDATA[E.]]></given-names>
</name>
<name>
<surname><![CDATA[Dras]]></surname>
<given-names><![CDATA[M.]]></given-names>
</name>
</person-group>
<source><![CDATA[Arabic dialect identification using a parallel multidialectal corpus]]></source>
<year>2015</year>
<conf-name><![CDATA[ International Conference of the Pacific Association for Computational Linguistics]]></conf-name>
<conf-loc> </conf-loc>
<page-range>35-53</page-range></nlm-citation>
</ref>
<ref id="B11">
<label>11</label><nlm-citation citation-type="confpro">
<person-group person-group-type="author">
<name>
<surname><![CDATA[Medhaffar]]></surname>
<given-names><![CDATA[S.]]></given-names>
</name>
<name>
<surname><![CDATA[Bougares]]></surname>
<given-names><![CDATA[F.]]></given-names>
</name>
<name>
<surname><![CDATA[Esteve]]></surname>
<given-names><![CDATA[Y.]]></given-names>
</name>
<name>
<surname><![CDATA[Hadrich-Belguith]]></surname>
<given-names><![CDATA[L.]]></given-names>
</name>
</person-group>
<source><![CDATA[Sentiment analysis of tunisian dialects: Linguistic ressources and experiments]]></source>
<year>2017</year>
<conf-name><![CDATA[ Third Arabic Natural Language Processing Workshop]]></conf-name>
<conf-loc> </conf-loc>
<page-range>55-61</page-range></nlm-citation>
</ref>
<ref id="B12">
<label>12</label><nlm-citation citation-type="confpro">
<person-group person-group-type="author">
<name>
<surname><![CDATA[Mulki]]></surname>
<given-names><![CDATA[H.]]></given-names>
</name>
<name>
<surname><![CDATA[Haddad]]></surname>
<given-names><![CDATA[H.]]></given-names>
</name>
<name>
<surname><![CDATA[Gridach]]></surname>
<given-names><![CDATA[M.]]></given-names>
</name>
</person-group>
<source><![CDATA[Polarity analysis of non figurative tweets: Tw-star participation on deft 2017]]></source>
<year>2017</year>
<conf-name><![CDATA[ 24e Conférence sur le Traitement Automatique des Langues Naturelles (TALN)]]></conf-name>
<conf-loc> </conf-loc>
<page-range>92-8</page-range></nlm-citation>
</ref>
<ref id="B13">
<label>13</label><nlm-citation citation-type="confpro">
<person-group person-group-type="author">
<name>
<surname><![CDATA[Mulki]]></surname>
<given-names><![CDATA[H.]]></given-names>
</name>
<name>
<surname><![CDATA[Haddad]]></surname>
<given-names><![CDATA[H.]]></given-names>
</name>
<name>
<surname><![CDATA[Gridach]]></surname>
<given-names><![CDATA[M.]]></given-names>
</name>
<name>
<surname><![CDATA[Babao&#287;lu]]></surname>
<given-names><![CDATA[I.]]></given-names>
</name>
</person-group>
<source><![CDATA[Tw-star at semeval-2017 task 4: Sentiment classification of arabic tweets]]></source>
<year>2017</year>
<conf-name><![CDATA[ 11th International Workshop on Semantic Evaluation (SemEval-2017)]]></conf-name>
<conf-loc> </conf-loc>
<page-range>664-9</page-range></nlm-citation>
</ref>
<ref id="B14">
<label>14</label><nlm-citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname><![CDATA[Piryani]]></surname>
<given-names><![CDATA[R.]]></given-names>
</name>
<name>
<surname><![CDATA[Madhavi]]></surname>
<given-names><![CDATA[D.]]></given-names>
</name>
<name>
<surname><![CDATA[Singh]]></surname>
<given-names><![CDATA[V. K.]]></given-names>
</name>
</person-group>
<article-title xml:lang=""><![CDATA[Analytical mapping of opinion mining and sentiment analysis research during 2000-2015]]></article-title>
<source><![CDATA[Information Processing &amp; Management]]></source>
<year>2017</year>
<volume>53</volume>
<numero>1</numero>
<issue>1</issue>
<page-range>122-50</page-range></nlm-citation>
</ref>
<ref id="B15">
<label>15</label><nlm-citation citation-type="confpro">
<person-group person-group-type="author">
<name>
<surname><![CDATA[Rosenthal]]></surname>
<given-names><![CDATA[S.]]></given-names>
</name>
<name>
<surname><![CDATA[Farra]]></surname>
<given-names><![CDATA[N.]]></given-names>
</name>
<name>
<surname><![CDATA[Nakov]]></surname>
<given-names><![CDATA[P.]]></given-names>
</name>
</person-group>
<source><![CDATA[Semeval-2017 task 4: Sentiment analysis in twitter]]></source>
<year>2017</year>
<conf-name><![CDATA[ 11th International Workshop on Semantic Evaluation (SemEval-2017)]]></conf-name>
<conf-loc> </conf-loc>
<page-range>502-18</page-range></nlm-citation>
</ref>
<ref id="B16">
<label>16</label><nlm-citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname><![CDATA[Rushdi-Saleh]]></surname>
<given-names><![CDATA[M.]]></given-names>
</name>
<name>
<surname><![CDATA[Martín-Valdivia]]></surname>
<given-names><![CDATA[M. T.]]></given-names>
</name>
<name>
<surname><![CDATA[Ureña-López]]></surname>
<given-names><![CDATA[L. A.]]></given-names>
</name>
<name>
<surname><![CDATA[Perea-Ortega]]></surname>
<given-names><![CDATA[J. M.]]></given-names>
</name>
</person-group>
<article-title xml:lang=""><![CDATA[Oca: Opinion corpus for arabic]]></article-title>
<source><![CDATA[Journal of the Association for Information Science and Technology]]></source>
<year>2011</year>
<volume>62</volume>
<numero>10</numero>
<issue>10</issue>
<page-range>2045-54</page-range></nlm-citation>
</ref>
<ref id="B17">
<label>17</label><nlm-citation citation-type="confpro">
<person-group person-group-type="author">
<name>
<surname><![CDATA[Samih]]></surname>
<given-names><![CDATA[Y.]]></given-names>
</name>
<name>
<surname><![CDATA[Attia]]></surname>
<given-names><![CDATA[M.]]></given-names>
</name>
<name>
<surname><![CDATA[Eldesouki]]></surname>
<given-names><![CDATA[M.]]></given-names>
</name>
<name>
<surname><![CDATA[Abdelali]]></surname>
<given-names><![CDATA[A.]]></given-names>
</name>
<name>
<surname><![CDATA[Mubarak]]></surname>
<given-names><![CDATA[H.]]></given-names>
</name>
<name>
<surname><![CDATA[Kallmeyer]]></surname>
<given-names><![CDATA[L.]]></given-names>
</name>
<name>
<surname><![CDATA[Darwish]]></surname>
<given-names><![CDATA[K.]]></given-names>
</name>
</person-group>
<source><![CDATA[A neural architecture for dialectal arabic segmentation]]></source>
<year>2017</year>
<conf-name><![CDATA[ Third Arabic Natural Language Processing Workshop]]></conf-name>
<conf-loc> </conf-loc>
<page-range>46-54</page-range></nlm-citation>
</ref>
<ref id="B18">
<label>18</label><nlm-citation citation-type="">
<person-group person-group-type="author">
<name>
<surname><![CDATA[Sayadi]]></surname>
<given-names><![CDATA[K.]]></given-names>
</name>
<name>
<surname><![CDATA[Liwicki]]></surname>
<given-names><![CDATA[M.]]></given-names>
</name>
<name>
<surname><![CDATA[Ingold]]></surname>
<given-names><![CDATA[R.]]></given-names>
</name>
<name>
<surname><![CDATA[Bui]]></surname>
<given-names><![CDATA[M.]]></given-names>
</name>
</person-group>
<source><![CDATA[Tunisian dialect and modern standard arabic dataset for sentiment analysis: Tunisian election context]]></source>
<year>2016</year>
</nlm-citation>
</ref>
<ref id="B19">
<label>19</label><nlm-citation citation-type="confpro">
<person-group person-group-type="author">
<name>
<surname><![CDATA[Yasavur]]></surname>
<given-names><![CDATA[U.]]></given-names>
</name>
<name>
<surname><![CDATA[Travieso]]></surname>
<given-names><![CDATA[J.]]></given-names>
</name>
<name>
<surname><![CDATA[Lisetti]]></surname>
<given-names><![CDATA[C. L.]]></given-names>
</name>
<name>
<surname><![CDATA[Rishe]]></surname>
<given-names><![CDATA[N. D.]]></given-names>
</name>
</person-group>
<source><![CDATA[Sentiment analysis using dependency trees and named-entities]]></source>
<year>2014</year>
<conf-name><![CDATA[ Twenty-Seventh International Florida Artificial Intelligence Research Society Conference]]></conf-name>
<conf-loc> </conf-loc>
<page-range>134-9</page-range></nlm-citation>
</ref>
</ref-list>
</back>
</article>
