Gains:
- Ability to recognize hallucination forms in physics (fake constant, fake law, unit error, wrong result, fake attribution) and catch them with the appropriate method.
- Ability to protect unpublished, confidential and personal data by anonymization and minimization by not entering them into unauthorized tools
- Being able to understand that the ultimate responsibility lies with humans by reporting the data honestly, without embellishing it, and transparently documenting the use of artificial intelligence.
Throughout this module, we covered the same discipline in each unit: artificial intelligence (AI) accelerates, humans verify. In this final unit, we bring together this entire verification practice under the umbrella of scientific integrity, the most fundamental value of physical science. Physics is the pursuit of producing verifiable knowledge about nature; At the heart of this endeavor are integrity, repeatability and responsibility. AI can be used to reinforce these values — but when used incorrectly, it threatens them all. In this unit, we will discuss the physical forms of hallucination, data and privacy responsibility, ethical boundaries, and why the ultimate responsibility always remains with humans.
Forms of hallucination in physics
Hallucination—AI's confident production of false information—takes unique and dangerous forms in physics. Recognizing these is the first step to catching:
type of hallucination
example
Capture method
Fitting fixed/data
An incorrect material density
Confirmation with reliable reference
False formula/law
A non-existent "principle" name
Textbook/SymPy derivation
Unit/order error
mistaking eV for J
Dimensional analysis, unit control
Incorrect numerical result
Improper code output
Run the code
apocryphal attribution
non-existent article
Verification with DOI
Wrongly reported result
Distortion of real article
Read the original source
The common thread is this: a large language model is designed not to tell the truth, but to produce "possible-looking text". In physics, the difference between "seemingly possible" and "true" is precisely the reason for science's existence.
Data, privacy and intellectual property liability
Physicists often work with sensitive data: unpublished experimental data, pre-patent designs, confidential measurements from industrial collaborations, research involving personal health or safety data. Entering such data into a general AI tool could make it accessible to third parties and possibly the model's training data—which would mean both a breach of confidentiality and a loss of scientific priority (the right to be the first to make a discovery). The rule is clear: do not enter unpublished, confidential or personal data into an AI tool that your organization has not approved. If you are going to enter, anonymize, minimize (give only what is necessary) and read the tool's data conditions.
Caution: Entering a laboratory's unpublished measurements or a company's confidential design parameters into an AI tool "for quick analysis" can be an irreversible mistake. Once data is shared, it cannot be recalled. If in doubt, do not enter data; Clarify your institution's policy and the privacy terms of the tool first.
Ethical boundaries and scientific integrity
A use may be technically possible and even legal, but still be unethical. There are several red lines in physics that concern scientific integrity:
Data fabrication or embellishment. Asking the AI to “generate points that better fit the yield” or “make the graph look more convincing” is data fraud. Data is reported as is; Outliers are eliminated only with a documented, legitimate reason.
Transparency of contribution. If you've used AI in a study — code writing, language correction, analysis — it's fair to make that clear at the appropriate place. Many journals and institutions require this.
Authorship and responsibility. AI is not a writer; cannot be responsible for the content of a work. The responsibility always lies with human writers. An error produced by AI cannot be defended by saying "the vehicle made it".
Reproducibility. A scientific result requires the method (including code used, parameters, AI prompts) to be clearly documented so that others can replicate it.
three mini cases
Case 1 — Confidential data leak. One researcher pasted measurement data from a non-disclosure agreement with a company into a generic AI tool “for a quick graph.” He later realized that this tool could store inputs. A breach of contract and possible loss of priority has occurred. Lesson: confidential data is never entered into an unapproved tool.
Case 2 — Beautification print. One student asked the AI for an edit that would "make the data look closer to the theory" because the experimental data did not fully match the theory. His advisor noticed this and stopped him: a discrepancy meant either a systematic error in the experiment or interesting new physics—something to be investigated, not hidden. Data were reported honestly and systematic error was found.
Case 3 — Gained transparency. One team wrote the data analysis code for a paper with the help of AI and clearly stated this in the methods section; shared the code and prompts as additional files. A referee reviewed the code and suggested a minor fix. Transparency caught the error before publication and increased the credibility of the study.
Four copyable templates
1) Hallucination screening:
Scan the following physics output for hallucinations: are there any bogus constants, bogus formula/law names, unit/order errors, unverified numeric results, and unconfirmed attributions? Flag each suspicious item and specify how to verify it (reference, code, DOI). Output: [here]
2) Data privacy pre-check:
Evaluate the following data/text before entering it into an AI tool: does it contain unpublished, confidential, personal or contractually protected information? If so, suggest how I can anonymise or minimize it. If unsure, say "check corporate policy". Content: [here]
3) Draft transparency/contribution statement:
In one study, I used AI for: [coding / language correction /analysis, etc.]. Write a short draft "AI use statement" that honestly documents this. It should also indicate which parts have been verified by humans.
4) Ensuring scientific integrity:
Perform an integrity check for the following analysis/result: data reported as is (no embellishment), are uncertainties reported, is the method documented in a reproducible manner, is the use of AI transparent? List any deficiencies. Content: [here]
Weak prompt / Strong prompt
Weak: "Edit the yield to better fit the theory and make the graph convincing."
Result: Data fraud; Direct violation of scientific integrity, irreversible loss of reputation.
Strong: "The yield does not fully agree with the theory. Help me analyze possible causes of this discrepancy (systematic error, statistical fluctuation, missing model, or a new effect); suggest how to report the discrepancy honestly without changing the data."
Result: An honest, investigative and faithful approach to scientific integrity; Conflict becomes an opportunity for discovery.
Common mistakes
- Confusing hallucination with self-confidence. Just because the AI speaks confidently is not a guarantee of accuracy; Even the most precise sentence may be fabricated.
- Entering confidential data into the vehicle without approval. Once shared, data cannot be retrieved; Priority and confidentiality are lost.
- Beautifying data. Changing data to "fit theory" is one of the most serious scientific crimes.
- Hiding the use of AI. Transparency is a requirement of honesty and is mandatory in many institutions.
- Putting the responsibility on the vehicle. “The AI did it” is not a defense; The ultimate responsibility always lies with the human being.
Tip: At the end of every AI-powered study, ask yourself three questions: (1) Has every fact, number, and attribution in this output been independently verified? (2) Did the data I used contain confidential/personal information that I should not enter? (3) Have I documented my AI usage and methodology transparently enough for someone else to replicate it? If you cannot answer a clear "yes" to these three questions, the study is not yet completed.
In summary
Scientific integrity is the basis of existence of physics and is even more important in the age of AI. Hallucination appears in physics in the form of false constants, false laws, unit errors, false conclusions and false references; each is captured with its own unique verification method. Confidential and personal data are not entered into unauthorized tools; data is never beautified; AI use is transparently documented; and ultimate responsibility always lies with man. There is only one discipline you learn throughout this module: AI is a powerful accelerator, but verification, integrity and responsibility rest on the shoulders of the competent physicist — that is, you. An unverified AI output is never a substitute for a physicist's confirmation.
Application task
Select an AI output (an account, an analysis, or a text) that you produced in this module. 1. screen for hallucinations with template: are there any dummy constants, spurious formulas, unit errors, or unverified attributions? With template 2, evaluate whether the data you are using is privacy-friendly. Finally, draft a short AI usage/transparency statement with template 3. Write down in 5-6 sentences: what did the scan reveal, what verification steps did you take?
checklist
- [ ] I independently verified every constant, formula, and result in the output.
- [ ] I have verified each citation with the DOI/source.
- [ ] I have checked that I have not entered confidential/personal data into the tool without approval.
- [ ] I reported the data as it is, without embellishment.
- [ ] I have documented my use of AI in a transparent and reproducible manner.
- [ ] I accepted that the ultimate responsibility was mine; I did not attribute any fault to the vehicle.
Module Exam
1. Which of the following is the most accurate positioning for artificial intelligence in physics?
- A) Artificial intelligence is a code, draft and editing assistant; Responsibility for unit, rank, verification and ethical decisions lies with the human being ✔
- B) Artificial intelligence is a calculator and the numerical results it gives are always accurate
- C) Artificial intelligence only works in calculation, it has nothing to do with experiment analysis and teaching
- D) Since artificial intelligence is more impartial than humans, unit and result decisions should be left to it.
Description: Artificial intelligence; It is an assistant who writes code, produces drafts, creates and organizes derivation frameworks. However, the ultimate responsibility for unit and order verification of the result, ensuring it with conservation laws, constant and reference confirmation belongs to the competent physicist. The big language model is not a calculator or physics engine; is a text generator that produces the 'next most likely word' and its unverified output may lead to an incorrect value, design or result.
2. What is the most accurate approximation for the numerical result of a computational code you print into a large language model?
- A) Trusting the number the model says 'this code gives that', because he wrote the code
- B) Run the code yourself and see the output; ✔ not relying on the model's verbal numerical prediction
- C) Accepting the code as correct just by looking at how it is read, without running it at all
- D) If the result is directly asked to the model again and it tells the same number, it is considered correct.
Explanation: An LLM may 'calculate' the output of the code he wrote in his head and return a wrong number; whereas when the same code is run it gives the correct result. What is deterministic is the code that is executed; It is not a verbal prediction of the model. Therefore, the code should always be run in your own environment and the output should be seen.
3. What is the most powerful way to test whether the simulation of a frictionless dynamic system accurately reflects physics?
- A) Seeing that the chart looks visually neat and convincing
- B) Running the simulation with only a single time step and trusting the result
- C) Monitoring a quantity that needs to be preserved, such as total energy or momentum, and checking whether it remains constant ✔
- D) Evaluating the simulation by how fast it runs
Explanation: In a frictionless system, total energy (and momentum, if any) must be conserved. Monitoring this magnitude throughout the simulation and seeing whether it remains constant reveals numerical flaws (e.g. the primitive Euler method constantly increasing the energy). If energy is drifting, the simulation is unreliable; Just because it produces a nice graphic doesn't mean it's correct.
4. Which of the following is the correct approach when fitting a model to experimental data?
- A) Choosing the higher degree polynomial that best fits the data because it passes through all points
- B) Giving the parameters as a single number without ambiguity
- C) Declaring the quality of fit simply by looking at how beautiful the curve looks.
- D) Selecting the model from physics, reporting the parameters with their uncertainties and evaluating the fit with the residuals ✔
Explanation: The model should be chosen from the physical law that the event must obey, not by looking at the data and wondering 'which one fits best'. With enough free parameters (e.g. a high degree polynomial) a 'perfect' but physically meaningless fit (overfit) can be made to any data. Parameters should be reported with their uncertainty, quality of fit should be evaluated with residuals, and physical plausibility should be checked.
5. Repeating a measurement many times reduces which type of uncertainty?
- A) Random uncertainty; systematic uncertainty does not decrease with repetition, the cause must be found and corrected ✔
- B) Systematic uncertainty; because it eliminates recalibration error
- C) Completely eliminates both ambiguities
- D) Does not reduce any of them; uncertainty is independent of the number of measurements
Explanation: Repeated measurement reduces random uncertainty (with standard error of the mean std/√N) resulting from random distribution of results. Systematic uncertainty, on the other hand, is an error that always deviates in the same direction (for example, an incorrectly calibrated balance) and does not decrease with repetition; However, the cause can be found and corrected. Reporting a 'very small' uncertainty with multiple measurements may hide underlying systematic error.
6. What is the most reliable independent way to verify an error propagation formula derived by AI?
- A) Relying on the formula to appear fluid and convincing
- B) Cross-checking with a Monte Carlo calculation that randomly samples the inputs with their uncertainty ✔
- C) Asking the same formula to artificial intelligence again and expecting it to give the same result
- D) Collecting the uncertainties directly and comparing them with the result
Explanation: AI often mixes partial derivative coefficients and square root structure in error propagation formulas. A powerful way to achieve this independently is Monte Carlo cross-checking: randomly sampling each input with its uncertainty, calculating the function thousands of times, and comparing the standard deviation of the output to the uncertainty given by the formula. The meeting of two independent methods increases confidence.
7. What is the correct division of labor between AI and SymPy in a long derivation of symbolic physics?
- A) SymPy generates the idea of the derivation, AI precisely verifies each step
- B) Both are unnecessary; long derivations must be done manually only
- C) AI builds the framework of derivation, SymPy precisely verifies each algebraic step ✔
- D) Artificial intelligence does the derivation on its own, there is no need for verification because it looks fluent
Explanation: AI strategizes the derivation and explains the path, but in long algebra it makes sign errors, escaped terms, and incorrect simplifications. SymPy, on the other hand, works according to rules and precisely verifies each algebraic step. The most powerful workflow combines the two: get the skeleton from the AI, source each step with SymPy, test the result with the derivative-integral inverse and the boundary case.
8. What does dimensional analysis tell about the accuracy of a physical formula?
- A) Dimensional analysis always proves definitively that a formula is correct
- B) Dimensional analysis has no practical value in physics
- C) Dimensional analysis only gives the constant factor of the formula, not its consistency
- D) A formula that does not hold dimensionally is absolutely wrong; dimensional analysis catches error without counting numbers ✔
Explanation: The two sides of a physical equation must be dimensionally equal before they can be numerically equal. A formula that does not hold dimensionally is simply wrong; this is captured from the count calculation. Dimensional analysis does not conclusively prove that a formula is correct (it may miss a dimensionless constant factor or angle), but it conclusively proves that it is incorrect; so it is the first and cheapest audit of every account.
9. What is the most critical first step when calculating the kinetic energy of an object traveling at 36 km/h?
- A) Converting all inputs into a single system of units (SI); Converting 36 km/h to 10 m/s ✔
- B) Putting the speed in km/h directly into the formula without converting it to m/s
- C) Multiplying only numbers without taking into account the units
- D) Assuming that there is no need to check the order after calculating the result
Explanation: SI based formulas (E = ½mv²) expect speed in m/s. Confusing km/h with m or hours with seconds in a calculation is one of the most common mistakes in physics. First all inputs must be converted into a single system of units (SI): 36 km/h = 10 m/s. Since AI can make silent errors in compound unit conversions, this conversion must be verified manually or with a unit-aware tool.
10. Which practice is correct in terms of 'honesty' in a scientific graph?
- A) Hiding uncertainty because error bars clutter the chart
- B) Label the axes with units, select the scale without distortion, and show uncertainty with error bars ✔
- C) Always start the y-axis in the range that will show the most dramatic difference
- D) Leaving the axes without units because the reader can guess the unit
Description: An honest scientific chart; it labels axes with magnitude and unit, chooses scale consciously and without distortion (starting the y-axis away from zero for exaggeration is misleading), and makes measurement uncertainty visible with error bars. Hiding uncertainty makes the data seem certain and creates a false impression of bias or agreement. The more convincing a graphic is, the more important the integrity check is.
11. What is the biggest risk and the correct discipline when getting literature and citation support from artificial intelligence?
- A) Artificial intelligence wastes time because it produces citations very slowly
- B) Sources suggested by artificial intelligence are always real, there is no need for verification
- C) Artificial intelligence can produce attributions that appear real but do not exist; Each citation must be verified with DOI/source and the literature must be found from the actual database ✔
- D) Made-up citations are easily recognized because they always use unreal journal names
Explanation: A large language model is not connected to an actual database; By mimicking citation patterns, it can produce articles that bear the real journal name, credible authors, and a reasonable year but are entirely non-existent. Therefore, literature should be found from real databases, AI should be used only to summarize/translate the actual text at hand, and each citation should be independently verified by DOI or title. An unverified citation is a violation of academic integrity.
12. What does it take to get reliable results when having AI summarize an article?
- A) Just giving the title of the article is enough, the model knows the content
- B) Using the summary directly without checking it at all, because it looks fluid
- C) Asking the model for the same summary twice and accepting it as correct if it is similar.
- D) Giving the original text to the model and comparing the produced summary with the original text to check that nothing has been added/distorted ✔
Explanation: Simply giving the AI the title of an article and asking for a summary will cause the model to come up with a 'plausible summary' without ever seeing the article; This summary may be the opposite of the truth. A reliable summary can only be obtained when the original text is fed to the model and the generated summary is compared with the original text. Additionally, it should be checked that the summary does not add results or numbers that are not in the text.
13. Which is true for the output of artificial intelligence in physics course material production?
- A) Each solution should be independently verified, the limits of analogies should be specified, and distractors should target real misconceptions ✔
- B) The answer key produced by artificial intelligence is always correct, there is no need for auditing
- C) Analogies are perfect in all cases, there is no need to specify their limits.
- D) Distractors should be chosen completely randomly, because the nature of the incorrect choice is unimportant.
Description: AI is fast at generating questions, examples, and analogies, but checking is essential on three points: the solution to each problem must be verified independently (by hand or with code), because the model may make sign/unit errors in the solution; Each analogy must be stated where it breaks down, otherwise it will give rise to a new error; Distractors should target true misconceptions. The responsibility lies with the teacher; An incorrect answer key leaves students with misunderstandings that are difficult to correct.
14. What is the behavior consistent with scientific integrity when experimental data does not fully agree with the theory?
- A) Having artificial intelligence organize the data to fit the theory and show the graph convincingly
- B) Investigating the reasons for the discrepancy (systematic error, lack of model, new effect) and reporting honestly without changing the data ✔
- C) Ignoring disagreement and reporting only points that fit
- D) Deleting the data and fitting a new experiment result
Explanation: Asking AI to 'arrange the data to better fit the theory' is data fraud and a gross violation of scientific integrity. Disagreement is something to be investigated, not hidden: it may indicate a systematic error, a statistical fluctuation, a lack of pattern, or an interesting new effect. Data is reported as is, with uncertainty; The ultimate responsibility lies with the person and 'the tool did it' is not a defence.