<?xml version="1.0" encoding="UTF-8"?>
<!DOCTYPE ArticleSet PUBLIC "-//NLM//DTD PubMed 2.7//EN" "https://dtd.nlm.nih.gov/ncbi/pubmed/in/PubMed.dtd">
<ArticleSet>
<Article>
<Journal>
				<PublisherName>دانشگاه تهران</PublisherName>
				<JournalTitle>پژوهشهای زبانشناختی در زبانهای خارجی</JournalTitle>
				<Issn>2588-4123</Issn>
				<Volume>10</Volume>
				<Issue>3</Issue>
				<PubDate PubStatus="epublish">
					<Year>2020</Year>
					<Month>10</Month>
					<Day>22</Day>
				</PubDate>
			</Journal>
<ArticleTitle>Reliability of Human Translations’ Scores  Using Automated Translation Quality Evaluation Understudy Metrics</ArticleTitle>
<VernacularTitle>پایایی نمرات ترجمه‌های انسانی با استفاده از ابزارهای خودکار جانشین ارزیابی کیفیت ترجمه</VernacularTitle>
			<FirstPage>618</FirstPage>
			<LastPage>629</LastPage>
			<ELocationID EIdType="pii">78592</ELocationID>
			
<ELocationID EIdType="doi">10.22059/jflr.2020.309025.751</ELocationID>
			
			<Language>FA</Language>
<AuthorList>
<Author>
					<FirstName>سمیه</FirstName>
					<LastName>کرمی</LastName>
<Affiliation>دانشجوی دکتری رشته ترجمه، دانشگاه اصفهان،
اصفهان، ایران</Affiliation>

</Author>
<Author>
					<FirstName>داریوش</FirstName>
					<LastName>نژادانصاری</LastName>
<Affiliation>استادیار گروه زبان و ادبیات انگلیسی، رشته آموزش زبان انگلیسی، دانشگاه اصفهان،
اصفهان، ایران</Affiliation>

</Author>
<Author>
					<FirstName>اکبر</FirstName>
					<LastName>حسابی</LastName>
<Affiliation>استادیار گروه زبان و ادبیات انگلیسی، رشته زبان‌شناسی همگانی، دانشگاه اصفهان،
اصفهان، ایران</Affiliation>

</Author>
</AuthorList>
				<PublicationType>Journal Article</PublicationType>
			<History>
				<PubDate PubStatus="received">
					<Year>2020</Year>
					<Month>09</Month>
					<Day>05</Day>
				</PubDate>
			</History>
		<Abstract>Considering the costly nature of translation quality assessment in terms of time, money and energy, it seems logical to benefit from the modern technologies that are introduced in the field of machine translation (MT). Automated Translation Quality Evaluation Understudy Metrics (ATQEUMs) are one of these technologies that have revealed a promising application in assessing the MT output quality. This study, however, attempts to examine the reliability of the scores provided by the lexical ATQEUMs to human translated texts (i.e. the ones provided by 51 senior students of translator training programs in Iran) using 1, 2, …, 5 reference translations successively and separately. To this end, an empirical applied study is conducted following a quantitative approach to assess the reliability of the lexical ATQEUMs’ scores in comparison to the expert scorers’ scores. The higher the correlation between the sets of scores (in different stages of using 1, 2, …, 5 reference translations), the higher the reliability is interpreted to be. The results of the Pearson correlation coefficient analysis revealed that using 5 reference translations had led to the highest correlations in 37.80% of cases, which is more than the number for any other situation considered (i.e. using 4 reference translations (3.65%), 3 reference translations (10.97%), 2 reference translations (31.70%), and 1 reference translation (15.85%)). However, using 2 reference translations achieved the second position in having the highest correlations which contradicted the hypothesis that more reference translations would lead to higher correlations and reliability.</Abstract>
			<OtherAbstract Language="FA">&lt;span lang=&quot;FA&quot; style=&quot;font-size: 11.0pt; line-height: 75%; font-family: &#039;B Nazanin&#039;;&quot;&gt;با توجه به ماهیت فرایند ارزیابی ترجمه که از لحاظ زمان، انرژی و هزینه قابل تامل می‌باشد، بهره‌گیری از فن‌آوری‌های نوین در حوزه‌ ترجمه ماشینی منطقی به نظر می‌رسد. ابزارهای خودکار جانشین ارزیابی کیفیت ترجمه یکی از این فن‌آوری‌ها است که در حوزه ترجمه ماشینی کاربرد دارد. این پژوهش در صدد یافتن پاسخ این سؤال است که پایایی نمرات این ابزارها در سطح واژگان به ترجمه‌های انسانی (۵۱ دانشجوی سال آخر رشته ترجمه در ایران) با استفاده از ۱، ۲، ... تا ۵ ترجمه معیار به صورت مرحله به مرحله و جداگانه چه تغییری می‌کند. لذا پژوهشی تجربی و کاربردی با رویکردی کمی برای محاسبه میزان پایایی نمرات این ابزارها در مقایسه با میانگین نمرات ۵ ارزیاب‌ متخصص انجام شد. میزان رابطه همبستگی میان این دو مجموعه نمره (در حالت‌های مختلف استفاده از ۱، ۲، ... تا ۵ ترجمه معیار) به منزله پایایی نمرات ابزار خودکار تفسیر شده است. نتایج تحلیل آزمون همبستگی پیرسون نشان داد که استفاده از ۵ ترجمه معیار در ۸۰/۳۷ درصد موارد منجر به بالاترین میزان رابطه همبستگی شده است که بیشتر از هر حالت دیگر در این پژوهش است (۴ ترجمه معیار (۶۵/۳درصد)، ۳ ترجمه معیار (۱۰.۹۷ درصد)، ۲ ترجمه معیار (۷۰/۳۱ درصد) و۱ ترجمه معیار (۸۵/۱۵ درصد)). بنابراین، فرضیه پژوهش تایید می‌شود که استفاده از ترجمه‌های معیار بیشتر منجر به رابطه همبستگی بالاتر و پایایی بیشتر نمرات می‌شود. در عین حال، استفاده از ۲ ترجمه معیار جایگاه دوم را از نظر دستیابی به بالاترین میزان رابطه همبستگی دارد و فرضیه پژوهش را نقض می‌کند.&lt;/span&gt;</OtherAbstract>
		<ObjectList>
			<Object Type="keyword">
			<Param Name="value">ارزیابی کیفیت ترجمه</Param>
			</Object>
			<Object Type="keyword">
			<Param Name="value">ابزارهای خودکار جانشین ارزیابی کیفیت ترجمه</Param>
			</Object>
			<Object Type="keyword">
			<Param Name="value">نمره‌دهی خودکار</Param>
			</Object>
			<Object Type="keyword">
			<Param Name="value">پایایی</Param>
			</Object>
			<Object Type="keyword">
			<Param Name="value">ترجمه معیار</Param>
			</Object>
		</ObjectList>
<ArchiveCopySource DocType="pdf">https://jflr.ut.ac.ir/article_78592_a98cad38031e4e5ebb9b6691b3a7e2b1.pdf</ArchiveCopySource>
</Article>
</ArticleSet>
