单位 6 / 12

偏见与正义

收益:

  • 了解数据是人工智能偏见的主要来源
  • 认识到偏见可能在工作场所造成的风险
  • 实施可采取的减少偏见的措施

人工智能表现为中立、客观的机器;但事实要复杂得多。 AI systems learn human biases in the data they are trained on and sometimes magnify and reproduce them. Prejudice is the tendency of a system to systematically favor or exclude certain people or groups. In this unit, we will consider where prejudice comes from, what concrete risks it poses in the workplace, and what can be done against it.这不仅仅是“道德”问题; It is a practical business matter that carries legal, commercial and reputational risk when mismanaged.

偏见从何而来?

The information of a language model comes from the huge texts that humans produce. These texts contain all the knowledge of human society, as well as all its stereotypes, imbalances and historical inequalities. The model does not distinguish between "good" and "bad" patterns;它学习数据中的任何内容。

示例:如果某些职业在过去的文本中主要与性别相关(例如“护士”与女性,“工程师”与男性),则模型会学习这种关联。她在完成句子时可能倾向于选择男性代词,“一位工程师走进办公室,他……”该模型并无恶意;它仅仅反映了数据的不平衡。

因此,偏差主要是由于它是数据的镜像,而不是有意识的编程错误。将问题解释为“有人编写了糟糕的代码”是一种误导;问题在于世界的不平衡被带入数据并从数据带入模型。

Tangible Risks in the Workplace

面积

偏见风险

可能的伤害

招聘

了解历史数据的趋势并消除某些群体

法律诉讼、丧失能力

客户沟通

刻板或排他性语言

Loss of reputation, customer flight

信用/风险

延续过去的不平等

歧视、制裁

绩效评估

将数据偏差带入决策中

不公平的晋升/奖金

内容/营销

Make specific group invisible

品牌受损

三个迷你箱

案例 1——广告缩小了语言范围。 A company started printing job postings with AI.过了一段时间,人们注意到这些广告以“年轻、有活力”、“攻击性目标”等表述来吸引特定人群,从而减少了应用程序的多样性。 The team added a step that checks each posting for “inclusive language”;几个月内,申请人的多样性显着增加。

Case 2 — Hidden bias in the summary.一位经理让 AI 总结了 30 名员工的反馈。 The recap seemed to concentrate the negative comments on one particular team;而在原始数据中,情况是平衡的。 The model exaggerated several strong statements.当经理回去检查原始数据时,他避免了误解。

Case 3 — The inclusivity mandate worked.一个营销团队将活动文案中的说明标准化为“使用包容性的、非特定的性别/年龄/群体语言,并在最后列出你的假设。”文本中的公式化表达明显减少;团队不必一遍又一遍地手动纠正每个输出。

Why Is Bias Hard to Recognize?

偏见最阴险的地方在于它往往是看不见的。输出看起来流畅、专业、合理;只有当人们仔细观察或衡量其对不同群体的影响时,其中的趋势才会变得明显。 It is misleading to say "there is no bias here" by looking at a single example;偏见往往不是表现在个别例子中,而是表现在系统的总体趋势中。

注意:“我是中立的,因此我使用的工具也是中立的”的假设是危险的。该工具的中立性取决于训练数据及其检查方式,而不是您的意图。

弱提示/强提示

Weak prompt: Describe the ideal candidate for this position.

结果:模型可能依赖于刻板印象并将特定的轮廓呈现为“理想”。

强烈提示:仅列出与该职位真正相关的可衡量能力。 Do not use any characteristics that are not related to the job, such as gender, age, marital status, hometown.最后,在“我避免的假设”标题下写下您遗漏的内容。

结果:评估仅限于与工作相关的标准,排除不相关的特征。

减少偏差的模板

Check the text below for inclusive language.标记暗示性别、年龄、种族、残疾或特定群体的陈述,并为每个陈述建议一个中立的替代方案。文本:[粘贴到此处]

仅根据这些客观标准进行评估:[标准]。不要考虑除这些标准(姓名、性别、年龄)之外的任何信息。

在末尾的单独标题下列出您在生成的文本中所做的假设。 So I can see hidden biases.

通过三个不同读者群体的眼睛评估此内容,并标记任何可能冒犯/冒犯他们中的任何人的陈述。内容:[粘贴到此处]

常见错误

常见错误

  • Assuming "neutral" AI output.即使输出看起来是中性的,它也可能带有数据偏差; Active supervision is required.
  • 看一个例子就说“没有偏见”。偏见隐藏在制度的总体倾向之中。
  • 将高影响力的决策留给无人监督的人工智能。人的权威在招聘、晋升和信用等决策中至关重要。
  • Forgetting to demand inclusivity. If you don't explicitly request it in the prompt, the model reverts to the default patterns.
  • It means "The machine said it." The person who uses the decision is also responsible for the consequences.

Justice is Everyone's Responsibility

Fighting prejudice is not just the job of the technical team. Every employee who uses AI output has the responsibility to question whether that output is fair. “The machine told me” is not an excuse; The person making the decision is also responsible for the consequences.

综上所述

  • The primary source of AI bias is not malevolence, but existing imbalances and stereotypes in the training data.
  • 偏见; It creates real legal and reputational risk in areas such as recruitment, customer communications, finance and evaluation.
  • Bias is often invisible and embedded in the overall disposition of the system; Looking at a single example is misleading.
  • To reduce: be aware, limit decision to job-related criteria, explicitly ask for inclusivity, put human control on high-impact decisions.
  • “The machine told me” is not an excuse;公平是每个使用输出的人的责任。

应用任务

告诉人工智能“描述一个理想的[任何职业]工人”并标记在输出中做出了哪些(可能是不必要的)假设。然后重复上面的“与工作相关的标准”模板,并比较两个输出的公平性。

清单

  • [ ] 据我所知,偏差的主要来源是训练数据。
  • [ ] 我可以说出工作场所偏见可能带来的具体风险。
  • [ ] 我知道偏见通常是看不见的,仅看一个例子会产生误导。
  • [ ] 我可以在提示中明确要求包容性和与工作相关的标准。
  • [ ] 我需要对高影响力决策进行人工监督。